Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Best Network Monitoring Tools in 2026: Compare Top Platforms

Most network monitoring tools alert you that a device is down. The best ones help you determine whether the problem is your WAN circuit, your ISP, or your SaaS provider before your users file a ticket. Traditional network monitoring tools were built for static networks. You poll devices, check interface counters, and still can’t explain why users are complaining about latency.
Sponsored Post

CloudWatch Logs to S3: The Easy Way

Many organizations use Amazon CloudWatch to analyze log data, but find that restrictive CloudWatch log retention issues hold them back from effective troubleshooting and root-cause analysis. As a result, many companies may be looking for effective ways to export CloudWatch logs to S3 automatically. Let's look at some of the reasons why you might want to export CloudWatch logs to S3 in the first place, along with some Amazon-native and open-source tools to help you with the process.

Beyond polling: Why enterprises are exploring network telemetry

Polling has been the go-to approach for network monitoring for years, and it still plays an important role in keeping networks healthy. But as networks become more distributed, application-driven, and data-intensive, simply polling devices more often isn't always the most efficient way to gain deeper operational insights. That's where network telemetry comes in.

Notes from the Field: Understanding "Lost connection" LAS activations in Citrix Virtual Apps and Desktops

With the transition from file-based licensing to the License Activation Service now complete, many Citrix administrators are spending more time in the Citrix Cloud licensing portal. As organizations continue to operate and troubleshoot LAS-based Citrix Virtual Apps and Desktops environments, it becomes increasingly important to understand what the licensing dashboard is actually showing.

The hard part of AI root cause analysis is no longer the model

Every few weeks someone tells me root cause analysis is a solved problem now: pipe your telemetry into an LLM, let it tell you what broke. I wish it were that easy. After years on this, I think "can AI do RCA?" is the wrong question, because doing RCA with an LLM is really two separate jobs, and the answer is different for each. They break in completely different ways, so it's worth pulling them apart.

New Feature: Automatic Snapshots When Latency Spikes

We’ve released an exciting new Lightrun capability: set a duration threshold on your Tic & Toc or Method Duration metrics, and Lightrun will automatically capture a snapshot whenever execution exceeds it. It takes moments to configure, and gives engineers the runtime context they need to understand why unexpected slow executions are occurring.