Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Claude Code Monitoring at Scale: Gateways and Routing With OpenTelemetry

Chelsea and I recently wrote a guide on how we monitor Claude Code usage internally with Bindplane. TLDR; We remotely manage a Bindplane Distribution of the OpenTelemetry Collector (BDOT) that runs on every engineer's laptop. This setup is great, but it has one downside. Sending to Google Cloud Monitoring, Swarmia, and any other destination directly from an engineer’s laptop is limited to local processing. You can’t get the benefit of centralized routing and processing on a gateway.

Build an SRE Agent Harness for AIOps Without Context Blowout

An agent harness for AIOps is the runtime layer that coding agents like Claude Code were never built to provide: context isolation, decision traceability, and gated execution for tools that touch production. Aura is Mezmo's open-source (Apache 2.0) agent harness, purpose-built for operations work rather than software development.

Skylar Advisor Guided Walkthrough

Learn how Skylar Advisor helps IT operations teams move beyond monitoring to AI-driven operational intelligence. In this walkthrough, you'll see how Skylar Advisor helps operators investigate issues, identify meaningful operational risks, collaborate more effectively, and predict potential problems before they impact services. In this video you'll discover Skylar Advisors key features like: By combining Ask Skylar, investigations, advisories, and predictions, Skylar Advisor helps IT teams reduce noise, focus on what matters most, and proactively improve service reliability.

Business intelligence plugins for Grafana: A support update

In January, we announced that Grafana Labs had assumed maintenance of the business intelligence (BI) plugins created by Volkov Labs, and committed to a six-month maintenance period. Today, we’re sharing an update: we're extending our maintenance commitment through the end of 2026. As announced earlier this year, that commitment includes maintaining compatibility with recent Grafana releases while handling bug fixes, security updates, and community contributions on a best-effort basis.

When and what should I be logging?

This is a follow-up to Sergiy’s post Errors, traces, logs, metrics: when to reach for what. Modern observability platforms, like Sentry, give developers a lot of choice. For a given problem, should you use traces, profiles, metrics, logs? If you take away one thing from this post, I hope it’s this: when in doubt, start by adding a few targeted log lines.

ITSM Knowledge Management: How to Build a Knowledge Base Your Team Will Actually Use

How many times should your service desk solve the same problem before it becomes shared knowledge? A senior agent on a 14-person service desk we worked with last quarter had answered the same question four times in two days for four different employees. The solution was already documented but buried in a wiki nobody could find. That is exactly the gap ITSM knowledge management is designed to close.

10 Best Endpoint Management Software Tools in 2026

What makes one endpoint management tool better than another? Not the feature list. Almost every tool claims patching, asset tracking, and automation. What matters is whether it holds up across a few hundred machines, and how much time it hands back to your team. For most IT teams, a good tool needs to: We looked at 10 of the best endpoint management software tools for 2026. We read through G2 and Gartner Peer Insights ratings, checked vendor pricing pages, and went through user reviews.

What Is Packet Loss? Causes, Symptoms & How to Fix It

In this video, learn what packet loss is, why it happens, and how it silently impacts your network performance even when monitoring dashboards appear healthy. Discover the most common causes of packet loss, how it affects applications like video calls and web services, and why identifying the root cause quickly is critical for maintaining a reliable network.