San Francisco, CA, USA
2016
  |  By Mike Goldsmith
Your LLM application emits telemetry unlike anything else in your stack. Model calls, tool invocations, retrieval steps, and token usage arrive as spans whose attributes carry entire prompts and completions. That data is bulky, it's full of user content you may not be allowed to export, and depending on which instrumentation each service uses, the same fact can arrive under different attribute names.
  |  By Fred Hebert
Back in July, I wrote about how the Tenant team (the team behind Honeycomb Private Cloud (HPC)) has embraced the code review bottleneck to focus more of its work. One of the other challenges we have is that we're downstream of almost all the other teams at Honeycomb, meaning that we have to package up everyone's code and services, and how it gets provisioned! This is something impossible to handle through code review since there are so many engineers on other teams, and so few of us.
  |  By Rox Williams
What happens when a software engineer who has put their entire identity into being the resident expert of an obscure technology or language with ten years of experience and who knows the codebase like the back of their hand now has to compete with AI? Psychologically speaking, according to Dr. Cat Hicks, author of The Psychology of Software Teams and founder of Catharsis, that's called identity threat, and it's a surefire way to feel unsafe. We were honored Dr.
  |  By Dan Juengst
Earlier this year, we introduced more capabilities to support agents in production, a more chaotic, complex environment that requires a tremendous amount of context to understand. Unlike tools that capture shallow, pre-aggregated metrics, or cannot join a metric, trace, and log in one query, Honeycomb retains the context and connective tissue from telemetry data to build a nuanced view of production.
  |  By Dan Juengst
A few months ago, we launched Agent Timeline to close the gap between knowing an agent failed and understanding why. It took the tangled reality of a multi-agent, multi-trace workflow and rendered it as a single, readable conversation: every LLM call, tool invocation, handoff, and downstream system span laid out in the order they happened.
  |  By Mike Goldsmith
You're producing more trace data than you want to pay to store, so you sample. A fixed 1-in-100 rate cuts your bill, but it's blind. It keeps 1% of your errors, 1% of the requests to that rarely-hit route, and 1% of the health checks, all at the same rate. The noisy traffic you care about least dominates what you keep while the traces you need during an incident are the ones most likely to be gone.
  |  By Jodi Sloan
For teams building AI agents, the feedback loop should already be a familiar idea: watch how the agent behaves, find what needs improvement, ship a change, and measure the result. In theory, each turn builds on the last until the loop becomes a flywheel and your agent is getting more effective with each turn. In practice, many of us are still in reaction mode. A user reports something strange, costs spike, or an eval score drops.
  |  By Charity Majors
Welcome to the third and final part of our series on AI norms and values. Parts of this doc were extracted and published separately on substack; as a whole, they describe the principles we hold pertaining to technology and AI, and the ethical commitments we make to each other and our customers. We set out to write about AI, and ended up writing about ourselves. These documents are not meant to be aspirational ones; they are derived from how we do our work every day in honeycomb.
  |  By Nick Travaglini
As agentic AI workflows gain traction within organizations, those organizations are asking how to account for their behavior while keeping costs manageable. Some are sticking with the old three pillars of observability approach: take a measurement to create a metric, record output to a log, and track serial progress with a trace. Each of these is useful, but treating them as distinct formats from the start means paying for them distinctly too. Separate storage doesn't come cheap.
  |  By Ken Rimple
I'm investigating repeated errors in my e-commerce application, and I need to get enough context in a single Honeycomb query to piece the entire picture together. Each query returns events based on the event's WHERE clauses, but I want to know several things from outside of the event that recorded an error. Things like: Those attributes are all over the trace. That's going to make a single query tough, right? Wrong!
  |  By Honeycomb
Honeycomb AI Ecosystem shows you the performance and cost of every AI agent in one live view and drills straight to the conversation behind any number, so you can diagnose your whole fleet in minutes without switching tools or rebuilding a single query.
  |  By Honeycomb
Watch a 2-minute demo video of Honeycomb Anomaly Detection, part of the Honeycomb Intelligence suite.
  |  By Honeycomb
A few months ago, Darragh Curran, CTO at Fin (formerly Intercom) set a public goal to double productivity and nearly tripled it instead. They did so by pulling a few levers: AI writing code at scale, building an AI-driven PR review system, leveraging as a trust mechanism, and with leadership becoming more hands-on through the transition. Charity wanted to pick Darragh’s brain on the messy bits, not just the highlight reel, so she invited him to participate in our first episode of Leading With Observability.
  |  By Honeycomb
It started with a single log line taking up a massive amount of volume: 500 million emissions per hour. Pulling that thread led Emma and Steven into Slack's broader logging pipeline: 311 billion logs per day at 4.4M/sec peak, with no volume limits, no per-service attribution, and no feedback to the teams generating the noise.
  |  By Honeycomb
'The runbook lost. The trace is the documentation now.' In his O11yCon 2026 closing keynote, Corey Quinn of Duckbill Group makes the case that when your primary reader is an, not a person, are the only pillar built to survive.
  |  By Honeycomb
It started with a single log line taking up a massive amount of volume: 500 million emissions per hour. Pulling that thread led Emma and Steven into Slack's broader logging pipeline: 311 billion logs per day at 4.4M/sec peak, with no volume limits, no per-service attribution, and no feedback to the teams generating the noise.
  |  By Honeycomb
Stripe shares lessons from building an incident investigation agent, from context-window blowups to why the final 5% still needs a human. In this O11yCon 2026 talk, they dig into what it takes to go from 'it works' to 'it works reliably,' including how pointing agents at like Honeycomb's speeds up on-call investigations.
  |  By Honeycomb
In this session at O11yCon, Purvi Kanal, Jamie Danielson, and Martin Holman demoed the new Canvas. Canvas now understands OpenTelemetry GenAI semantic conventions, and can show agent invocations, LLM calls, and tool calls all in one trace view. Humans and agents can work in the same place, with skills that let each team encode their own expertise so both the agent and their colleagues can use it. Multiplayer support means you can see your teammates' cursors and share charts.
  |  By Honeycomb
At Slack, between 100 to 200 users per day use Honeycomb for client observability, tracing, instrumentation, analysis of performance, frontend issues, investigating incidents, or just looking into production issues.
  |  By Honeycomb
Watch Nathen Harvey's full talk at O11yCon 2026, Honeycomb's observability conference, and enjoy Christine Yen's intro as well.
  |  By Honeycomb
Honeycomb is an event-based observability tool, but you can-and should-use metrics alongside your events. Fortunately, Honeycomb can analyze both types of data at the same time. When maturing from metrics-based application monitoring to an observability-based development practice, there are considerations that can make the transformation easier for you and your team.
  |  By Honeycomb
Evaluating observability tools can be a daunting task when you're unfamiliar with key considerations and possibilities. This guide steps through various capabilities for observability tooling and why they matter.
  |  By Honeycomb
This document discusses the history, concept, goals, and approaches to achieving observability in today's software industry, with an eye to the future benefits and potential evolution of the software development practice as a whole.

Honeycomb is a tool for introspecting and interrogating your production systems. We can gather data from any source—from your clients (mobile, IoT, browsers), vendored software, or your own code. Single-node debugging tools miss crucial details in a world where infrastructure is dynamic and ephemeral. Honeycomb is a new type of tool, designed and evolved to meet the real needs of platforms, microservices, serverless apps, and complex systems.

Honeycomb provides full stack observability—designed for high cardinality data and collaborative problem solving, enabling engineers to deeply understand and debug production software together. Founded on the experience of debugging problems at the scale of millions of apps serving tens of millions of users, we empower every engineer to instrument and query the behavior of their system.