Operations | Monitoring | ITSM | DevOps | Cloud

How we cut Spark compute costs by 44% with agentic AI and Datadog Jobs Monitoring

Spark jobs only get more expensive and harder to debug as they scale. It’s a problem we’ve run into ourselves. Our Referential Data Platform team builds and maintains the knowledge graph that maps relationships between customers’ observability entities. ServiceQueryEdge is at the center of that graph, mapping service entities to their associated metric and log queries.

From Cleanup to Animation in One Workspace: Redefining the Editing Loop

For the past several months, I have been watching a pattern emerge in how people actually use AI image tools. The pattern is not about any single feature. It is about how often a task that starts as a simple cleanup request evolves into something entirely different. A user uploads a product shot to remove a stray reflection. Then they wonder what the same image would look like with a different background. Then they think about turning it into a short social video. Each step is logical, but traditional workflows treat each step as a separate job requiring a separate tool.

Anthropic's Mythos, Glasswing, and how the industry must move forward | Harness Blog

When Anthropic broke the news of Mythos and Project Glasswing, the security community did what it always does. It published a flurry of papers asking "What does this mean for security?" It's a reasonable instinct, but it's the wrong question. The real question is who actually owns the problem?

Monitor LLM routing with the Kubernetes Inference Extension

If you serve LLMs on Kubernetes without inference-aware routing, your load balancer is likely wasting inference capacity. Generic HTTP traffic management blindly routes requests, assuming the backends in your cluster are interchangeable. But your model-serving backends are stateful and unevenly prepared to handle any given request. As a result, requests are often routed to the backend that’s not the one best suited to respond.

Uber blew its annual AI budget in 4 months

Uber burned through its entire annual AI budget in under 4 months. Here's what went wrong — and what every engineering org should be doing instead. The data: 80% more code is getting pushed with AI… but only 18% of AI-written code actually ships to production. That's not a productivity story. That's a spend problem. If you're scaling AI tooling without real-time monitoring and guardrails, you're Uber.

Konstruct product updates: Global resources, MCP support, and smarter permissions

May has been one of our busiest months yet for Konstruct. Across three releases, 0.5, 0.5.1, and 0.5.2, we've shipped some of the most requested platform-level changes since we launched: a unified model for sharing resources across organizations, native support for AI-driven workflows via MCP, a completely redesigned API keys experience, and a cleanup to how permissions actually work in multi-org environments. Let's walk through what shipped and why it matters.