%term

Three roles you need for reliability success

May 7, 2024 By Gavin Cahill In Gremlin

It’s one thing to say that reliability is a priority for your organization, and a whole other thing to make actual, demonstrable improvements in the availability of your applications. Sadly, it’s common for organizations to invest time, money, and effort into improving reliability only to barely nudge the needle on incidents and downtime. But there are hundreds of companies successfully improving their reliability posture—and doing it at enterprise scale.

Read Post

Gremlin

Read more about Three roles you need for reliability success

How to build reliable services with unreliable dependencies

May 2, 2024 By Andre Newman In Gremlin

In an earlier blog, we looked at slow dependencies and how they can impact the reliability of other services. While we explored what happens when dependencies are degraded, what happens when dependencies outright fail? What can you do when your application or service sends a request to another service, and nothing comes back? We’ll answer this question by using Gremlin to proactively test a service with multiple dependencies.

Read Post

Gremlin

Read more about How to build reliable services with unreliable dependencies

Confident Cloud Migrations How a Top 5 Bank Ensured Reliability With AWS and Gremlin

Apr 29, 2024 By Gremlin In Gremlin

In today's competitive landscape, migrating to the cloud brings substantial benefits, but the cloud’s new architectures and tools also bring new reliability risks and considerations. The challenge: Enterprises have to figure out how to capitalize on the benefits of the cloud while ensuring a seamless, reliable transition. This webinar offers a look at how to provide application reliability before, during, and after migrations with AWS and Gremlin.

View Video

Gremlin

Read more about Confident Cloud Migrations How a Top 5 Bank Ensured Reliability With AWS and Gremlin

Building Resilience in the Cloud With the AWS Well Architected Framework and Gremlin

Apr 29, 2024 By Gremlin In Gremlin

Reliability and resilience in the cloud requires a different approach. Thankfully, the AWS Well-Architected Framework is a proven blueprint for cloud architects and engineering leaders seeking to design and operate resilient systems on AWS.

View Video

Gremlin

Read more about Building Resilience in the Cloud With the AWS Well Architected Framework and Gremlin

How to make your services resilient to slow dependencies

Apr 24, 2024 By Andre Newman In Gremlin

When discussing reliability, we tend to focus on the things that we have control over: applications, virtual machine instances, deployment patterns, etc. But this ignores a significant and ever-growing part of nearly all modern software: dependencies. Dependencies are services that provide extra functionality for other services and applications. For instance, many websites depend on databases, caches, payment processors, and similar services in order to function.

Read Post

Gremlin

Read more about How to make your services resilient to slow dependencies

Hitting reliability goals in the face of layoffs

Apr 23, 2024 By Jeff Nickoloff In Gremlin

It’s never easy when layoffs hit your organization. In addition to the personal impact of losing friends and coworkers from your team, those who remain are left trying to achieve the same business goals with less people and resources. Unfortunately, layoffs and restructuring have become a common part of business. But you’re not alone. Your partners (including Gremlin) are here to help you navigate your new reality.

Read Post

Gremlin

Read more about Hitting reliability goals in the face of layoffs

How to ensure your Kubernetes Pods and containers can restart automatically

Apr 16, 2024 By Andre Newman In Gremlin

As complex as Kubernetes is, much of it can be distilled to one simple question: how do we keep containers available for as long as possible? All of the various utilities, features, platform integrations, and observability tools surrounding Kubernetes tend to serve this one goal. Unfortunately, this also means there’s a lot of complexity and confusion surrounding this topic. After all, most people would agree that availability is important, but how exactly do you go about achieving it?

Read Post

Gremlin

Read more about How to ensure your Kubernetes Pods and containers can restart automatically

How to ensure your Kubernetes cluster can tolerate lost nodes

Apr 12, 2024 By Andre Newman In Gremlin

Redundancy is a core strength of Kubernetes. Whenever a component fails, such as a Pod or deployment, Kubernetes can usually automatically detect and replace it without any human intervention. This saves DevOps teams a ton of time and lets them focus on developing and deploying applications, rather than managing infrastructure.

Read Post

Gremlin

Read more about How to ensure your Kubernetes cluster can tolerate lost nodes

How to test your systems for scalability and redundancy with Fault Injection

Apr 11, 2024 By Gremlin In Gremlin

Part of the Gremlin Office Hours series: A monthly deep dive with Gremlin experts. Do you know if your services can tolerate losing a node? What about an entire availability zone? Or a region?‍ Large-scale outages aren’t unheard of. When you’re running critical services, it’s vital that those services can keep running even if an AZ or region fails. In addition to failing over, these services also need to scale quickly so traffic shifts don’t overwhelm your systems. How do you prove that a service is both scalable and redundant? The answer is with Fault Injection.

View Video

Gremlin

Read more about How to test your systems for scalability and redundancy with Fault Injection

How to standardize resiliency on Kubernetes

Apr 10, 2024 By Gavin Cahill In Gremlin

There’s more pressure than ever to deliver high-availability Kubernetes systems, but there’s a combination of organizational and technological hurdles that make this ‌easier said than done. Technologically, Kubernetes is complex and ephemeral, with deployments that span infrastructure, cluster, node, and pod layers. And like with any complex and ephemeral system, the large amount of constantly-changing parts opens the possibility for sudden, unexpected failures.

Read Post