Operations | Monitoring | ITSM | DevOps | Cloud

Cloud Outage Preparedness: On-Call Lessons for 2026

Cloud outage preparedness stopped being a nice-to-have this month. In a span of roughly 48 hours, Microsoft Azure lost a big chunk of its West US footprint and Amazon Web Services dropped connectivity between its us-west-2 region in Oregon and the Seattle metro. The AWS event alone rippled outward and knocked DoorDash, Reddit, Hulu, Apple Pay, Snapchat, Fortnite, and the PlayStation Network offline for millions of users, according to incident trackers. Neither outage was caused by anything exotic.

Tech Talk | From Insight to Action: Reducing Metrics Costs in Splunk Observability

Explore how to leverage observability insights to reduce metric costs with Span Observability. Led by Tomasz Romaniuk and Martyna Karbownik from Cisco, this tech talk covers the significance of metric time series in impacting cardinality and consequently driving metric costs. After a theoretical introduction, the presentation dives into practical real-life use cases demonstrating how to effectively utilize available tools for cost reduction. The session concludes with an overview of future developments and a Q&A segment for audience inquiries.

4 Cloud-Native Challenges AI SRE Is Solving in 2026 and the 3 New Ones to Look Out For

AI SRE is making real strides in resolving some of the greatest pains related to incident response, troubleshooting, and complex root cause analysis. The on-call rotation, the war room, the week-long RCA, and the ticket queue that ate a third of every platform engineer’s week all look different now than they did two years ago.

What's New in InfluxDB 3.11: A Significant Performance Upgrade for Complex Time Series Workloads

Summary InfluxDB 3.11 delivers major performance and data management updates for growing time series workloads, with significantly faster queries on recent data, expanded support for wide and ultra-sparse schemas, and a more predictable resource profile under load. InfluxDB 3 Enterprise also adds backup and restore, bulk Parquet import, row-level deletes, and a built-in Explorer UI.

Eliminating Digital Friction with Nexthink Spark Episode 2: Fixing an Unresponsive Browser

Browser performance issues can disrupt productivity and create unnecessary frustration for employees. In Episode 2 of Eliminating Digital Friction with Nexthink Spark, see how Spark quickly investigates an unresponsive browser, identifies the root cause, and helps IT resolve the issue faster. Discover how AI-powered investigations enable proactive IT and deliver a better digital employee experience.

Shipped: Compare Periods, side-by-side cost comparison in Explorer

When spend moves, the first question is always “compared to what?” Answering it used to mean pulling up two tabs with two different date ranges and toggling between them, or squinting at a single delta number without the ability to customize what you’re comparing against. Compare Periods puts both periods in front of you at once, so a spike, a regression, or a shift shows up on the chart and in the cost table, without you doing the math yourself.

AI cost monitoring: what it is, how it works, and why real-time visibility matters

AI cost monitoring is the continuous tracking of AI and LLM spend in real time, broken down by the models, features, teams, and customers generating it. It is not the same as reading the monthly bill - done well, it shows spend as it happens, flags anomalies before they become invoices, and connects every dollar to an outcome so finance can protect AI ROI instead of explaining it after the fact.

Agentic AI cost: why agents burn tokens and how to control it

Agentic AI cost is what you pay to run AI agents, and it is mostly tokens. An agent does not answer once. It loops, calls tools, reads the results, and reasons again, re-sending a growing context every step. Anthropic found agents use about 4x the tokens of a chat, and multi-agent systems about 15x. You control it by capping runs, right-sizing the architecture, routing, caching, and measuring cost per task, then tying every agent to the AI ROI it produces.