%term

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Introduction to Ingesting Logs into Loki with Fluentd and Fluent Bit | Zero to Hero: Loki | Grafana

Jul 19, 2024 By Grafana In Grafana

Have you just discovered Grafana Loki and plan to use FluentD or Fluent Bit as your telemetry collector? Or are you trying to decide which agent is right for you? In this "Zero to Hero" episode, we cover the basics of FluentD and Fluent Bit, highlighting their differences and helping you determine when to use one over the other. Additionally, we guide you through configuring both agents' Loki plugins to write logs directly into Loki.

View Video

Grafana

Read more about Introduction to Ingesting Logs into Loki with Fluentd and Fluent Bit | Zero to Hero: Loki | Grafana

Learning Moment: Effective Customer Communication During Incidents - Enhance Visibility & Response with Uptime.com

Jul 19, 2024 By Jonathan Franconi In uptime

The recent global outage caused by an operating system update reminded me of how vulnerable we are today and most importantly, how close we are always teetering on global scale incidents with millions of interconnected dependencies. When the base of the house collapses, everything built on top is impacted. Those of us in IT Operations, Monitoring, Observability (insert the current acronym), etc., know firsthand this risk; we face it every day.

Read Post

uptime

Read more about Learning Moment: Effective Customer Communication During Incidents - Enhance Visibility & Response with Uptime.com

Chaos Testing Explained

Jul 19, 2024 By Shanika Wickramasinghe In Splunk

Chaos testing is a part of site reliability engineering (SRE). In chaos testing, we intentionally break things in and around a given application, in order to: The purpose of chaos testing is to assess how software systems respond to scenarios like network outages, hardware failures, database failures, and server or cluster node failures in the infrastructure.

Read Post

Splunk

Read more about Chaos Testing Explained

Monitoring Healthtech Applications with Custom Metrics

Jul 19, 2024 By Lauren Barnes In MetricFire

Staying healthy is essential. Luckily, nowadays, tracking health and wellness is easier than ever. This article will discuss how monitoring allows developers to ensure that their health applications run smoothly so people can stay healthy.

Read Post

MetricFire

Read more about Monitoring Healthtech Applications with Custom Metrics

How OTel Empowers You to Handle Unified Data

Jul 19, 2024 By ObservIQ In ObservIQ

Discover the power of OpenTelemetry to consolidate your telemetry data. Our expert-led workshop demonstrates standardization techniques for metrics, logs, and traces. Delve into real-world applications, including capturing Prometheus metrics, managing logs with FluentD/Bit, and collecting traces with Jaeger.

View Video

ObservIQ

Read more about How OTel Empowers You to Handle Unified Data

Sumo Logic's Next-Gen Apps

Jul 19, 2024 By Sumo Logic In Sumo Logic

Sumo Logic's next generation apps introduce new features that were previously not available with the Classic Apps. This video will explain what Next-Gen apps are, why our customers are encouraged to prefer them over our classic apps. The video also demonstrates how to install a next-gen app and upgrade or uninstall the installed app.

View Video

Sumo Logic

Read more about Sumo Logic's Next-Gen Apps

July 19th global IT outage reminds us of digital complexity

Jul 19, 2024 By Dritan Suljoti In Catchpoint

As we write, on Friday July 19th, a massive global cyber outage is continuing to take down critical services around the world dependent on Microsoft-based computers.

Read Post

Catchpoint

Read more about July 19th global IT outage reminds us of digital complexity

Global Microsoft Outage and Preventing Future Vulnerabilities

Jul 19, 2024 By Mishal Alam In uptime

In a recent unexpected turn of events, a faulty component in the latest CrowdStrike Falcon update led to widespread outages, crashing Windows systems globally. The repercussions were felt across various sectors, including airports, TV stations, hospitals, and even emergency services in the U.S. and Canada. The glitch, affecting both Windows workstations and servers, resulted in massive outages, bringing entire companies to a standstill and crashing fleets of hundreds of thousands of computers.

Read Post

uptime

Read more about Global Microsoft Outage and Preventing Future Vulnerabilities

The IT Scramble is On with a Microsoft Outage: Incident MO821132 - July 18, 2024

Jul 19, 2024 By Sara Purdon In Martello Technologies

On July 18, 2024 at 6:38 pm ET, Vantage DX, Martello’s Microsoft 365 and Teams performance management solution, started to see indicators of a likely Microsoft outage impacting users’ ability to access various Microsoft 365 apps and services. Almost an hour later at 7:41 pm ET Microsoft issued a statement on X.

Read Post

Martello Technologies

Read more about The IT Scramble is On with a Microsoft Outage: Incident MO821132 - July 18, 2024

Understanding and Troubleshooting Out of Memory Error Code 137

Jul 19, 2024 By Dmitry Maximov In StackState

If you've encountered the dreaded "exit code 137" error message while working with Docker, Kubernetes, or other containerized environments, you're not alone. This error can be frustrating and difficult to troubleshoot, but understanding its causes and solutions can help you keep your applications running smoothly. This comprehensive guide will delve into the intricacies of error code 137, its common scenarios, and strategies to resolve it.

Read Post