Operations | Monitoring | ITSM | DevOps | Cloud

Grafana Cloud updates: onboard teams with new AI-powered tooling, secrets management for enhanced security, and more

We consistently roll out helpful updates and fun features in Grafana Cloud, our fully managed observability platform powered by the open source Grafana LGTM Stack (Loki for logs, Grafana for visualization, Tempo for traces, and Mimir for metrics). In case you missed them, here’s our monthly round-up of the latest and greatest Grafana Cloud updates. You can also read about all the features we add to Grafana Cloud in our What’s New in Grafana Cloud documentation.

Nginx Logs & Performance Monitoring with Loki and Telegraf | MetricFire

When a web service slows down or errors spike, metrics can tell you what changed (active connections rise, error rate increases), but the root cause can sometimes be found in your logs (which IPs are hammering POST endpoints, 4XX/5XX occurrences). Put the two together and you get the full observability picture. Time-series metric trends to spot incidents, and line-level details to fix them fast.

How our engineers use AI for coding (and where they refuse to)

Okay, picture this: if you drew a Venn diagram of folks in tech right now, it'd probably look something like this: You'll probably find yourself in one of those circles, right? I’m guilty of falling in the intersection! Because let's be real, the 'will AI replace developers by 20xx?' debate is everywhere – Reddit, Hacker News, team Slack and even your local cafe. Well, we decided to go straight to the source.

Datadog governance 101: From chaos to consistency

As your organization scales, managing observability resources and usage becomes increasingly important. More users and teams mean more dashboards, tags, API keys, and costs to manage. The job of keeping track of these resources and ensuring that they’re compliant can quickly grow in complexity.

How we saved $1.5 million per year with Cloud Cost Management

In collecting and analyzing trillions of events each day, Datadog ingests a massive amount of data. We spend substantially to process and store this data in the cloud, and teams across the organization are committed to optimizing the return on this investment. To this end, our FinOps analysts have always tracked the costs of delivering our services and identified opportunities for savings.

Incident post-mortems: the complete, blameless guide

Most companies run post-mortems like autopsies. They dissect the corpse, assign blame, and file it away. The body count keeps rising. Here's what actually works: post-mortems as learning machines. Systems thinking over finger-pointing. Patterns over pain. What you'll get: A copy-paste template, real metrics that matter, and the mindset shift that turns outages into intelligence. Who this is for: SRE leads tired of repeating incidents. Engineering managers who want learning over theater.

Why Firms Are Adopting Payment Orchestration Now

In the fast-paced digital economy, businesses of all sizes must prioritize efficient payment processing to meet rising customer expectations and accommodate a diverse array of payment methods. Payment orchestration has emerged as a transformative solution, streamlining transactions, enhancing security, and optimizing the overall payment experience.

Solana's New Crypto Phone Targets Global Market with $450 Price Point

Solana pursues the mission to make cryptocurrencies a part of everyday life with the introduction of a new, crypto-friendly smartphone: the Seeker Chapter 2. Whereas the first model - Saga 1 - was produced in limited quantities of 20,000 units, this new version is set for a far broader, international audience.

Spectrum Delivers Bare-Metal RPC Infrastructure for Next-Gen Blockchain Operations

In today's fast-evolving web3 environment, infrastructure plays a decisive role in how decentralized applications (dApps) perform and scale. Spectrum, a global Remote Procedure Call (RPC) provider, is meeting this challenge head-on with a bare-metal infrastructure that spans continents and supports over one billion daily RPC requests across more than 175 blockchain networks.

How to Set Up and Manage LTO Tape Backup Systems That Last

Building a dependable LTO tape backup system starts with more than just the right hardware. From physical setup to ongoing tape management, each step plays a direct role in long-term data protection. Skipping planning or using mismatched components can lead to wasted time, damaged media, and incomplete backups. Hardware, software, labelling, storage, and tape rotation: This guide has you covered. It's all here, explained simply. We'll also look at what's usually missed. Testing. Without routine validation, your backups can fail silently. Keep your system running smoothly!