Chaos Engineering

What are the four Golden Signals?

Sep 2, 2022 By Andre Newman In Gremlin

When it comes to building reliable and scalable software, few organizations have as much authority and expertise as Google. Their Site Reliability Engineering Handbook, first published in 2016, details their practices to maintain reliability as Google scaled. But when you have over a million servers running thousands of services across more than twenty data centers, how do you monitor them in a consistent, logical, and relevant way?

Read Post

Gremlin

Read more about What are the four Golden Signals?

Four tests to measure and improve reliability: what matters and how it works

Sep 2, 2022 By Andre Newman In Gremlin

Legendary race car driver Carroll Smith once said, "until we have established reliability, there is no sense at all in wasting time trying to make the thing go faster." Even though he was referring to cars, the same goes for technology: no amount of code optimization or new features can replace stable systems. Unfortunately, much like race cars, it's hard to know that a system is unreliable until it blows a tire, the brakes stop working, or the steering wheel comes off the column.

Read Post

Gremlin

Read more about Four tests to measure and improve reliability: what matters and how it works

How to add a Golden Signal to a service in Gremlin RM

Sep 2, 2022 By Gremlin In Gremlin

In this video, we show you how to add a Golden Signal to a service. Gremlin uses your Golden Signals to ensure your services are still healthy and responsive during reliability tests. You can configure Golden Signals to use an existing monitor in your observability tools, such as Datadog, New Relic, or Prometheus. We recommend adding all four Golden Signals to each of your services to ensure comprehensive coverage.

View Video

Gremlin

Read more about How to add a Golden Signal to a service in Gremlin RM

How to add a Service to Gremlin Reliability Management (RM)

Sep 2, 2022 By Gremlin In Gremlin

This short demo video shows you how to add a Kubernetes service to Gremlin Reliability Management (RM). We'll walk you through selecting the parts of your infrastructure that make up your service, identifying processes for dependency detection, and adding your Golden Signals.

View Video

Gremlin

Read more about How to add a Service to Gremlin Reliability Management (RM)

Introduction to Gremlin Reliability Management (RM)

Sep 2, 2022 By Gremlin In Gremlin

Gremlin Reliability Management helps teams standardize and automate reliability, one service at a time. In this video, we walk through the platform by showing you how to add your services to Gremlin, integrate your Golden Signals, run reliability tests, and generate reliability scores.

View Video

Gremlin

Read more about Introduction to Gremlin Reliability Management (RM)

What is Chaos Engineering? A Guide on Its History, Key Principles, and Benefits

Aug 18, 2022 By Joey D'Antoni In SolarWinds

Many organizations invest in high availability and disaster recovery for their key applications. Too many of these organizations, however, forego the most important aspect of this process—testing the failover process regularly. Whether gripped by the fear of downtime or dreaded DNS problems, development teams are frequently hesitant to test out what they’ve built in the real world.

Read Post

SolarWinds

Read more about What is Chaos Engineering? A Guide on Its History, Key Principles, and Benefits

#DevOpsSpeakeasy at #KubeCon EU 2022 with Julie Gund on Chaos Engineering

Aug 17, 2022 By JFrog In JFrog

Julie Gund from Gremlin speaks about Chaos Engineering - the practice of proactively injecting failures into your systems to find vulnerabilities before your customers experience them, ultimately reducing chaos.

View Video

JFrog

Read more about #DevOpsSpeakeasy at #KubeCon EU 2022 with Julie Gund on Chaos Engineering

Chaos Engineering: What Is It & How Does It Work?

Aug 17, 2022 By Noor-ul-Anam Ruqayya In Blameless

Distributed software systems have many points of failure. Can the process of chaos engineering help identify problems and gauge resiliency?

Read Post

Blameless

Read more about Chaos Engineering: What Is It & How Does It Work?

Why SREs Need to Embrace Chaos Engineering

Jul 20, 2022 By xMatters In xMatters

Reliability and chaos might seem like opposite ideas. But, as Netflix learned in 2010, introducing a bit of chaos—and carefully measuring the results of that chaos—can be a great recipe for reliability. Although most software is created in a tightly controlled environment and carefully tested before release, the production environment is harsher and much less controlled.

Read Post

xMatters

Read more about Why SREs Need to Embrace Chaos Engineering

How to define and measure the reliability of a service

Jul 14, 2022 By Andre Newman In Gremlin

More and more teams are moving away from monolithic applications and towards microservice-based architectures. As part of this transition, development teams are taking more direct ownership over their applications, including their deployment and operation in production. A major challenge these teams face isn't in getting their code into production (we have containers to thank for that), but in making sure their services are reliable.

Read Post