Your On-Call Rotation Has a Single Point of Failure, and It Gets the Flu Every Winter
Most teams design on-call for the failures they can see in a dashboard. A region goes down, a deploy goes sideways, a certificate expires at 2 a.m. The rotation exists so that someone is always there to catch it. Far fewer teams design for the failure that takes out the catcher: the on-call engineer wakes up with a fever, and the plan for that is usually a Slack message and hope.
Treat a sick engineer the way you would treat any other dependency outage. It is predictable, it is seasonal, and it has a blast radius that grows with every shortcut in the rotation design.
Sick engineers are an availability problem
Flu season in the United States runs through fall and winter, and the CDC reports that activity most often peaks between December and February. That is the same window as holiday code freezes, year-end traffic peaks, and the thinnest staffing of the year. The question is not whether someone on the rotation will be out during a peak. It is how many weeks of the quarter the rotation is running on one person.
The math is unforgiving on small teams. Google's SRE book puts the minimum size of a single-site on-call team at eight engineers, so that each person carries primary or secondary duty one week a month and on-call stays under 25 percent of their time. Many teams run with three or four. On a four-person rotation, one illness means the secondary becomes the primary with no secondary behind them. Two illnesses, which is how respiratory viruses move through an office, and the rotation is a single engineer for a week.
So count depth, not names. For each shift, ask who the backup is, who the backup's backup is, and whether either of them has production access and current context. If the honest answer is "the person who is sick," you have found the single point of failure.
Presenteeism is an incident risk
The instinct of a conscientious engineer who gets sick during their shift is to tough it out. From a reliability standpoint, that is the wrong call. Fever, poor sleep and cold medication degrade exactly the capacities that incident response depends on: reading a graph correctly, remembering which runbook applies, noticing that the rollback command is pointed at the wrong environment. Nobody would accept a deploy from a build agent that was running at half capacity. The pager deserves the same standard.
There is also a containment issue. The CDC's guidance for respiratory illness is to stay home until symptoms are improving and there has been no fever for at least 24 hours without fever-reducing medication, and to take extra precautions for five days after that. An on-call engineer who comes into the office to be close to the war room is a vector into the rest of the rotation.
The fix is cultural before it is technical. Make "I am sick, handing over" an expected message, not a confession. The team lead's job is to acknowledge it, reassign the shift, and never ask the sick engineer to stay available "just in case." If the lead cannot do that without the rotation collapsing, the rotation is too shallow, and that is the lead's problem to fix, not the engineer's.
Remove the friction that keeps sick people online
Some of the pressure to work through an illness comes from the rotation itself. Some of it comes from HR policy that was written for a different kind of job. A sick-day rule that requires same-day medical documentation sends a feverish engineer to an urgent care waiting room for several hours to collect a piece of paper, and the predictable result is that they skip the visit and stay logged in instead.
Policy can take that pressure off in two ways. First, align the documentation threshold with how short illnesses actually run. Most acute respiratory infections resolve within a few days, and a note requirement that starts on day one mostly produces presenteeism. Second, accept documentation that does not require the engineer to leave the house. For the short absences this covers, getting a doctors note online from a licensed physician who reviews the case the same day is now routine, and the note carries a license number and a verification code that HR can check in seconds. The engineer rests, the rotation adjusts, and nobody spends a sick day in a waiting room.
None of that replaces a real medical visit when someone is genuinely unwell for longer, and the policy should say so. The point is to stop treating a two-day cold as a paperwork event.
The sick-coverage playbook
Write this down before the first cold front. Everything below fits on one page in the on-call handbook.
Minimum depth. Set a floor for how many qualified engineers must be available for each rotation week, and treat a breach of the floor the way you treat a breached SLO. If the team cannot reach the floor, merge rotations, borrow from a sister team, or reduce the on-call surface until it can.
A named sick-coverage path. Primary is sick, so the secondary takes primary. Who takes secondary? Decide in advance, by shift, and publish it in the schedule tool, not in someone's head. The escalation path should end at a manager who has agreed, in writing, to carry the pager if the chain runs out.
A fifteen-minute handover template. A sick engineer should be able to hand over in a quarter of an hour, by text if their voice is gone. The template has four lines: open incidents and their current state, anything deployed or changed in the last 24 hours, scheduled changes in the next 48 hours, and anything the engineer was watching that has not paged yet. Keep it in the incident channel so the next person finds it without asking.
Access that works without the sick person. The backup must already have production access, the paging app installed, the VPN configured, and the runbooks bookmarked. Discovering a missing permission at 3 a.m. while the only person who can grant it is asleep with a fever is a self-inflicted outage.
A quarterly drill. Once a quarter, pick a shift and pretend the primary is out. Have the backup actually take the pager for a day. Every gap this exposes is a gap a real illness would have exposed in December, at a worse time, with less notice.
Recovery before return. Returning engineers rejoin as secondary first. A day of shadowing restores context and keeps a still-foggy brain away from the production console.
The reliability case
Teams spend real money on redundancy for hardware that fails a few times a year. The humans on the rotation fail more often than that, on a schedule the CDC publishes annually, and the cost of ignoring it is paid in incident minutes and in the engineers who learn that getting sick is a problem for them to absorb alone. Design the rotation for the winter you know is coming, make handing over the pager the normal thing to do, and clear the paperwork out of the way of people who should be asleep.