Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

How savepoints quietly throttled our Postgres queue

At incident.io we are huge fans of Postgres; we've written about it a lot over the years, including how to choose the right indexes and how we're proud of being boring (The Pet Shop Boys). We use Postgres as our primary transactional database, which as of today has ~900 tables, and counting! The vast majority of our codebase does something along the following lines: read some data from Postgres, execute some business logic, then write that data back to Postgres. It is not, however, always that simple.

Ship faster, improve reliability, and control CI costs with Datadog CI/CD Optimization

AI-assisted development can increase the rate at which teams produce code, but teams only realize those velocity gains if CI can keep pace. More pull requests (PRs) mean more builds, tests, and pipeline executions. Slow jobs leave developers and coding agents waiting for feedback, flaky failures consume time in reruns and investigations, and unnecessary test execution increases runner demand as delivery volume grows.

Shipped: See what your AI spend is actually paying for

Most AI spend comes in with no tags and no owner attached. Your provider console shows total spend, maybe broken out by API key or model. It won’t tell you that the sales team spent $1,700 on Claude this week, let alone what the work was. And the problem is growing. McKinsey found that 56% of organizations now use AI in three or more business functions. More teams means more spend, and most companies respond with a spending cap. Set it too low and you slow down the work you wanted AI to help with.

Canada Data Center Development: Measuring Responsible AI Growth

Western Canada is becoming a live test of whether sovereign AI capacity can be built responsibly at scale, and the answer will depend less on what operators promise than on what they can measure and show. Meta’s planned C$13-billion Alberta data center, BCE’s expansion of its Saskatchewan project to a 1.2-GW hub, and the federal Responsible Data Centre Development Principles all point to the same requirement: operational transparency that regulators, utilities, and communities can verify.

Run your first workflow in minutes, no sales call

The regression nobody catches passes a busy review and ships. An off-by-one, a change that reads as sensible and quietly breaks something, gets a nod from a tired reviewer and lands in production, where it erodes trust one small defect at a time. You can have an AI code reviewer running on your own repository in the time it takes to read this page. Get started without having to book a demo or contact sales.

AI Agent Context Explained: What Agents Can't See in Your Infrastructure

The "C" word is a controversial subject in the US, but we have to talk about "context", and what it means to an AI Agent. For starters, agents can only act on what's in their context window. Everything outside of it is a guess. In application code that limit is usually an annoyance.

Why We Built the Komodor Agentic Operations Platform: Q&A with CEO Ben Ofiri

Komodor spent years building an AI SRE platform before the category had a name. With the launch of the Komodor Agentic Operations Platform, it’s opening that engine up so enterprises can build, run, govern and optimize their own agents in production. Following the launch, co-founder and CEO Ben Ofiri sat down to talk about why now is the right time for agentic operations, what breaks between prototype and production, and where operations will head next.

Beyond Traditional Observability: Turning Technical Insight into Operational Intelligence

Observability has become a central part of modern IT operations and for good reason. Metrics, logs and traces give technical teams detailed evidence about how applications, infrastructure and services are behaving. Such evidence helps them investigate performance degradation, identify abnormal behavior and understand what changed around the time an issue occurred.

Building a Self-Service Knowledge Base Employees Want to Use

Most organizations already have policy documents, troubleshooting guides, knowledge articles, resolved tickets, runbooks, and internal wikis. They're not starting at zero, but despite that investment, employees continue to open tickets for questions the organization has already answered. The problem is not always a lack of knowledge. More often, employees cannot find the right information quickly enough to trust self-service as their first option.