Operations | Monitoring | ITSM | DevOps | Cloud

How to verify your Azure Application Gateway is zone-redundant

Having a redundant failsafe is one of the best things you can do to ensure high availability in the cloud. It’s rare for cloud regions to go offline, but it can happen, even on major platforms like Azure. While you might not have control over your provider’s reliability, you do have control over your own, and redundancy is a key part of that. Here’s how you can check if your Application Gateways are availability zone (AZ) redundancy, and how to verify redundancy using active testing.

Managing Kubernetes node drains with Pod Disruption Budgets

Kubernetes normally excels at preserving uptime during maintenance tasks, but it’s not always perfect. Even something as benign as consolidating nodes after a traffic spike could take your application offline if not done carefully. This is where PodDisruptionBudgets (PDBs) come in. In this blog, we’ll explain why PDBs are important, what the risks are of not implementing them, and how you can find out which of your own deployments are missing PDB definitions. ‍

Gremlin app for Dynatrace - DEMO!

Dynatrace gives engineering teams deep, real-time visibility into every service they run. That visibility is the foundation of every effective reliability practice, and it's exactly the foundation Gremlin is built to extend. Once you can see how your distributed systems behave today, the next step is knowing how they'll behave under failure tomorrow—and to do it before those failures happen.

Safer Kubernetes rollouts with minReadySeconds

Picture the scene: you’ve just deployed a rolling update to your service. Half of your pods are running the new version, they all passed their readiness checks, and Kubernetes terminated the old replicas. Suddenly, the new pods start throwing 503 errors. Thankfully, you still have pods on the old version, so you stop the update. If the rollout had been a little bit faster, you’d have an outage. This is the failure mode minReadySeconds exists to prevent.

Optimizing Kubernetes pod deployments for reliability with topology spread constraints

If you’re like many Kubernetes users, you don’t pay much attention to where or how Kubernetes distributes your pods. As long as they’re running, it doesn’t matter where they get deployed, right? Surely Kubernetes will use some complex algorithm to figure out the most reliable way to distribute your pods across the cluster…right? Pod distribution plays a much bigger role in reliability than you might think.

What your AI SRE can't see (and what you can do about it)

AI SRE is having a moment. The category pulled in massive funding rounds over the last two years, Gartner published its first market guide, and vendors are promising everything from 90% faster resolution to fully autonomous incident response. If you run an engineering organization, someone has probably pitched you an AI SRE in the last quarter. And let’s be honest: faster triage, less alert fatigue, and automated frontline response are wins for understaffed teams.

Managing slow container starts with Kubernetes readiness probes

Imagine if your workday started as soon as you woke up. Before you can even start your coffee maker, email alerts are flooding in, coworkers are pinging you on Slack, and your phone is buzzing nonstop with reminders. You haven’t even pulled the covers back, and your boss is asking you about deliverables. This is what Kubernetes pods deal with every day. Unless, that is, you use readiness probes.

The Gremlin app for Dynatrace: resilience testing and reliability scoring, built on the observability you already trust

Dynatrace gives engineering teams deep, real-time visibility into every service they run. That visibility is the foundation of every effective reliability practice, and it's exactly the foundation Gremlin is built to extend. Once you can see how your distributed systems behave today, the next step is knowing how they'll behave under failure tomorrow—and to do it before those failures happen.

Eliminate Reliability Blind Spots in AWS, Azure, and GCP

Cloud resilience often feels like an uphill battle. When you’re overseeing hundreds of applications across different providers, identifying potential failure points manually is nearly impossible. You’re left trying to find the needle in a haystack—a needle that could take down your entire application at any moment. To truly protect your uptime, you have to break the cycle of reactive troubleshooting.

More Resilience, Less Overhead: How to Modernize Disaster Recovery Testing

• Disaster recovery planning is essential for ensuring digital services remain online in the face of catastrophic failures or outages. When a major digital infrastructure outage occurs, systems need to be set up to automatically respond and restore functionality as quickly as possible. But no matter how in-depth your disaster recovery plan is, it’s still only theoretical until it’s thoroughly tested under realistic failure conditions, which is why testing is often mandated by leadership and regulators.