London, UK
2021
  |  By Engineering
At incident.io we are huge fans of Postgres; we've written about it a lot over the years, including how to choose the right indexes and how we're proud of being boring (The Pet Shop Boys). We use Postgres as our primary transactional database, which as of today has ~900 tables, and counting! The vast majority of our codebase does something along the following lines: read some data from Postgres, execute some business logic, then write that data back to Postgres. It is not, however, always that simple.
  |  By Article
At incident.io, we've spent the last two years building Investigations, our AI SRE. When you get paged, it starts investigating straight away, looking across your telemetry, recent deploys, past incidents, docs and code, and posts what it's found in your incident channel (or on your phone, if it's 2am and you're still deciding whether you need to get out of bed). By the time you open your laptop, you're starting at step six of triage rather than step one.
  |  By Engineering
Six weeks into my internship, I was handed the biggest project I'd ever worked on: WhatsApp notifications for on-call paging. It was bigger by a large margin. When we scoped it out it broke into about a dozen chunks, each roughly the size of a whole project I'd done before. This is what I learned from it, and what it was like leading a project of that size as one of the most junior engineers at the company.
  |  By Engineering
Like many modern software stacks, the incident.io platform is predominantly event-driven. For example, whenever you send us an alert, post a message to our agent on Slack, or update an entry in your Catalog - these are all events that then get enqueued on a message topic, meaning any of our downstream components that are interested in that event can subscribe and react asynchronously, such as sending a push notification or posting a reply to you in Slack.
  |  By Data
We’ve previously written about how deeply embedded data is in people’s day-to-day work at incident.io, and I’d have it no other way — demand for data is undoubtedly a good thing. What risks breaking at scale, however, is everything downstream of that demand: data-team capacity gets stretched thin, dashboard sprawl outpaces anyone's ability to maintain it, and stakeholders can't reach an answer without going through the data team.
  |  By Data
Historically, non-technical stakeholders would’ve had most of their data questions answered either through pre-built dashboards or by asking their Data team (or equivalent). Self-serve analytics tools went a step further by offering safe, governed datasets built by Data teams which let non-technical users dig into data without having to worry about how it joins together, how metrics like “revenue” are defined, and so on.
  |  By Engineering
As the size and complexity of their relational database workload grows, every company eventually goes through the process of off-loading work on a read replica. It comes with lots of benefits, but at a cost of increased complexity. This article is about how we dealt with that, a lot of learnings, and some useful techniques. incident.io is an incident management product relied on by thousands of customers to be the thing that supports them through anything from a minor blip to a full outage.
  |  By Article
Our On-call product has a lot of great features: configuring escalation paths, viewing rotas and schedules, requesting cover, etc. However, when framing its reliability, we reduce it down to two critical pieces of functionality: It’s not that we’re happy if only these parts are working, but they are the most important parts. In this post, I'll go into more detail on how we think about their reliability.
  |  By Article
There's a conversation I keep having with our design partners at incident.io. It starts when I ask "what are you doing with AI internally?" and lands in a similar place every time. The shape of how their engineering teams work is changing fast. Not in vague "AI is transforming everything" ways, but in concrete, repeatable patterns. Different companies are building the same things. The frontier teams are six to twelve months ahead of the average, and they're describing the same future.
  |  By Article
When thinking about Service Level Objectives (SLOs) and contractual Service Level Agreements (SLAs) for availability, I always like to put the percentages into concrete numbers. It’s easy to lose track of what’s meant when saying “99.95%” availability, and even more is lost when thinking how much harder it is to achieve 99.99% compared to 99.95%. On a monthly basis, and in concrete terms, 99.95% availability means you get 21 minutes and 55 seconds of downtime.
  |  By incident-io
Zendesk replaced 15 years of homegrown incident tooling and PagerDuty by migrating 1,200 engineers across 150 teams onto incident.io in just 10 weeks, cutting mean time to triage by 32%, saving $500k+ in year one, and eliminating 800+ hours of annual toil, with zero incidents on go-live day. Tom Monaghan (VP of Engineering Productivity & Product Reliability) and Anna Roussanova (Engineering Manager) share how they pulled it off and what's next as Zendesk helps build Investigations, our AI agent that starts digging into incidents the moment an alert fires.
  |  By incident-io
We're announcing the PagerDuty Rescue Program. PagerDuty worked. For a long time, it was the standard. But the world's changed, and PagerDuty hasn't. The single biggest reason teams stay on PagerDuty isn’t the product - it’s the pain of leaving. So, we’ve removed every barrier. You've wanted out for a while. Now, nothing is stopping you.
  |  By incident-io
We rebuilt our post-mortems from the ground up. In this video, Pete and the engineering team talk through how they built it: the decisions they made, the problems they were solving, and what it took to ship AI-native post-mortems.
  |  By incident-io
OpsGenie is going away in 2027, forcing a migration decision for thousands of teams. But this isn't just a tooling swap — it's a rare chance to upgrade how you respond to incidents. Because the real pain in incident response isn’t paging. It’s everything that happens after the alert: coordination, clarity, communication, ownership, and follow-through. Most teams solve this through heroics and tool-juggling across chat, tickets, and docs. That approach doesn't scale.
  |  By incident-io
A full walkthrough of our completely rebuilt post-mortems experience. We cover AI-generated first drafts from your incident data, accuracy review, inline rewriting, a collaborative editor with live incident context, meeting notes with Scribe, and management tooling including dashboards, exports, and analytics. Post-mortems are included in incident.io Response. AI features and Scribe are available on Pro and Enterprise plans.
  |  By incident-io
Everyone your sales team is reaching out to is drowning in emails. The way to cut through isn't to send more of them. It's to get personal, get creative, and get bold. That's the philosophy baked into incident.io's sales culture: experiment constantly, celebrate the inputs as much as the wins, and never play it safe. This video gives you a real look at what it's like to be part of a sales team at one of the most exciting startups right now. There are many more wins to come, and we want the right people here for them.
  |  By incident-io
When an incident hits, every second counts. The response team at incident.io builds the tools that make sure engineers aren't flying blind when it matters most. Sam, Tech Lead of the response team, takes us inside what it's really like to build the core of incident.io: the high technical bar, the art of prioritisation, and why there's no shortage of meaningful work to do. If you're an engineer who wants to work on something that genuinely makes other engineers' lives better, this one's for you.
  |  By incident-io
Working on AI in incident management means there's no playbook. No million blogs. Just building at the forefront of what's possible with AI models.In this video, Martha, Product Engineer on our AI team, talks about what it's really like working with AI that helps engineers respond to incidents faster. This covers the shift from traditional engineering, learning the personalities of different AI models, and why you need to embrace constant change when new models drop all the time.
  |  By incident-io
Post-mortems are required, time-consuming, and widely disliked — but they’re also one of the biggest opportunities to improve reliability. In this webinar, we talked about how to run post-mortems that actually lead to learning and improvement. This covered why most post-mortems fall flat, how to structure them effectively, and walk through a real example to show what good looks like in practice. The goal: fewer wasted hours, better outcomes, and post-mortems that actually matter.

Create, manage and resolve incidents directly in Slack. Leave the admin and reporting to us.

Improving your incident response, visibility, and ability to learn:

  • Less faffing, more fixing: We take care of the admin during incidents, so you can save your brainpower for the decisions that matter.
  • Divide and conquer: We make sure everyone’s role is clear, track who’s working on what, and help you escalate if you need extra help.
  • Get up to speed, at speed: Get everyone on the same page from the moment they join the incident, and help stakeholders stay in the loop.
  • Timelines, in no time: Constructing an incident timeline for review is important, but time consuming. We’ll build one for you in real-time, and keep it constantly up to date.
  • Data and insights you can trust: You’ve already paid for your incidents. By surfacing the data you need to make decisions, we help you get your money’s worth.

Incident response for your whole organisation.