Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

Fiber Paths and Failsafes: Why Your Network Design Matters

Redundancy isn’t just a buzzword – it’s the design principle keeping modern AI and cloud applications online. In this Uplink episode, Kevin Schlosser, Interconnection Product Manager at NTT Global Data Centers, explains how resilient infrastructure is engineered to expect failure but remain operational. We explore: Diverse entry points and fiber path management AI-driven bandwidth growth: 100G standard, 400G emerging Cooling innovations for intense compute workloads Why providers without their own fiber may offer the most resilient paths.

Mike Long and DORA Community Discussion - Software Delivery Governance

Manual governance in regulated industries is like steering a ship with last year’s map. Approvals, ticket queues, and after-the-fact evidence collection slow delivery and increase risk. By the time an audit arrives, teams are scrambling to prove they followed the process. Watch Kosli’s Mike join Nathen Harvey at DORA to unpack why this happens — and what continuous, automated governance can do to fix it.

Cortex MCP set up

Learn how to set up the Cortex MCP in under 5 minutes. The MCP integrates directly into your IDE, giving instant access to Cortex data without leaving your coding environment. It reduces context switching by enabling natural questions about services and teams, and streamlines workflows with real-time data from Cortex, Jira, GitHub, and more.

Using Claude to power up your onboarding

I joined incident.io about ten weeks ago, having been in my previous role for four and a half years. Being a new starter was an unusual feeling for me, and there's been a huge amount to learn; but by lunch on my second day (!) I had started shipping value to our customers. A large part of hitting the ground running has been having a colleague alongside me, who I can pester with questions, who doesn’t get offended when I write in all capitals, and often praises me for being absolutely right!

Zero-downtime deployment with Flagsmith and CircleCI

As developers, we continually strive to improve our software. This often means rolling out new software features at a rapid pace. However, deploying new features to production is not without risk. From no real production testing to limited rollback options, traditional deployment can quickly become frustrating. The worst issues, though, usually stem from one thing: buggy features making their way into the hands of users.
Sponsored Post

Traffic Replay: Production Without Production Risk

The software and product life cycle is fraught with pitfalls and tradeoffs. While testing applications under production-like load is critical to ensuring the reliability, performance, and security of your data storage and software services, you need to do this testing without actually affecting the production data and systems. In essence, you have to pull off the impossible - be as close to production as you can without actually being production.

Stop Trying To Cut Cloud Costs, Start Trying To Price AI Correctly

Most SaaS companies aren’t spending too much on AI. They’re just completely screwing up how they price it. You feel the budget pressure. The OpenAI and Anthropic bills keep climbing. Finance is starting to twitch. So the instinct is to cut. Trim back experiments. Cap usage. Beg your team to “optimize.” You can’t cost-cut your way out of a pricing failure though. And most of the time, that’s all this is — a pricing failure.

How HireVue Turned Cloud Cost Chaos Into A Competitive Edge

When you’re a global leader in AI-assisted hiring, speed matters. Not just in matching candidates to jobs, but in making the engineering and financial decisions that keep your platform running efficiently. For HireVue, fragmented infrastructure, manual processes, and sprawling spreadsheets turned cloud cost management into a time-consuming spelunking expedition.