Operations | Monitoring | ITSM | DevOps | Cloud

Kubernetes AI SRE Agent Finds a Crash Loop Nobody Asked About: AURA

You ask for a routine health check and expect a clean baseline. What came back was a pod that had restarted 788 times, unrelated to the question. AURA is connected to a Kubernetes cluster and to Prometheus through read-only MCP servers, running as one coordinator with two specialized workers. The prompt is one sentence: check the health of the cluster, and confirm whether all the pods are running. What comes back is not a baseline. AURA names the state as CrashLoopBackOff and attaches the restart count to it.

No Custom Adapter: AI SRE Agent AURA Debugs Product Catalog in Dash0

The platform shows you which service is failing and which paths it touches, and stops there. Point AURA at the same telemetry and the cause comes back too. Dash0 shows the product catalog service in a failed state across the selected window, with errors on the path from the frontend service.

KPI cards: build a reliability dashboard that doesn't force tradeoffs

This week's Feature Friday: Principal Product Manager Christine Byun walks through KPI cards, a new way to build custom dashboards in Engineering Intelligence. KPI cards pull key metrics, like change failure rate and rollback frequency, into compact tiles so they stay visible without taking up chart space. That means the metric you're actively working, incidents, in this demo, gets full-size room, without losing sight of the rest of your system.

Data Center: More or Less | SolarWinds TechPod

In this episode, Sean and Crystal explore the complex and rapidly evolving landscape of AI, data center impacts, regulation challenges, and societal implications. They discuss the urgency of establishing standards and the lessons from historical industrial revolutions to navigate AI's future responsibly.

Monitoring Oracle ASM with Custom Metrics | The Tony and Tonie Show Ep 49

Even small Oracle ASM issues can become big database problems. Here's how to spot the warning signs early. Tony and Tonie discuss how Redgate Monitor custom metrics help teams close a common monitoring gap: surfacing Oracle ASM health and performance issues before storage pressure, rebalancing problems, or disk group failures become database incidents.

Creating Escalation and Regular Groups in OnPage

Learn how to create and configure an Escalation Group in OnPage with this step-by-step how-to guide. This video walks through how to create an escalation group, a regular group, configure escalation intervals and factors, enable Round Robin, set failover OPIDs, and add a Fail Report email address. With escalation groups, OnPage can route critical alerts to team members in a predefined order and automatically move to the next responder when needed—helping ensure time-sensitive notifications don’t go unanswered.

How to Create and Import Contacts in the NEW OnPage Web Console

A step-by-step guide for OnPage’s new web management console, including how to create a single contact and how to create multiple contacts at once by importing an Excel spreadsheet. Feel free to comment below with any questions! Whether you’re in IT or healthcare, OnPage helps teams manage critical alerting and communication to ensure urgent messages reach the right people at the right time. If you’re not yet using OnPage or want to see how it works, request a demo or speak with a member of our team to learn more.