The Failure Mode Your Runbook Probably Does Not Cover

Operations teams rehearse plenty of scenarios. Failed deployments, database corruption, certificate expiry, a region going dark, the on-call engineer who cannot be reached. What gets rehearsed far less often is the building losing power for eleven hours, because that feels like somebody else's problem, filed under facilities alongside the air conditioning and the parking barrier. It stops being somebody else's problem at the moment the UPS batteries drain and everything still running on premises goes down at once.

The Gap Between Ride-Through and Actual Backup

An uninterruptible power supply is not a backup power system, and the distinction matters more than the terminology suggests. A UPS exists to bridge the seconds or minutes between utility failure and either a generator picking up load or a controlled shutdown completing. Sizing one for extended runtime is possible and rapidly becomes expensive and impractical. A generator is what turns a brief interruption into an outage the business can sit through, and organizations that installed a UPS and considered the problem solved discover the difference during their first prolonged event. Specialist suppliers such as Oregon Generators exist because sizing, installing and maintaining standby generation is a discipline in its own right, and it is one that most IT teams have no reason to have developed internally.

Cloud Migration Did Not Remove the Problem

There is a comfortable assumption that moving workloads to a hyperscaler makes power somebody else's concern. It relocates a portion of the concern and leaves a surprising amount behind. Network termination equipment, firewalls, switches and the office connectivity your staff depend on all still sit in your building. So does anything that stayed on premises for latency, licensing, regulatory or hardware reasons, which in most organizations is more than the architecture diagram suggests. A team with everything in the cloud and no power in the office is still a team unable to work, which makes this an availability question rather than purely an infrastructure one.

Test It or Do Not Claim It

Standby generation has a specific failure pattern: it works perfectly during monthly no-load tests and fails during the actual event. Fuel degrades over time and needs treatment or rotation. Batteries that start the generator age quietly. Automatic transfer switches seize if they never operate. Cooling and fuel systems develop faults that only appear under sustained load. The only meaningful test is a periodic load test where the generator actually carries real demand for an extended period, and the only meaningful documentation is a record of those tests. An untested generator is a plan rather than a capability. Fuel duration deserves its own examination. A tank sized for eight hours is adequate for a typical utility fault and inadequate for a regional weather event, and the assumption that a delivery can be arranged mid-outage is exactly the assumption that fails when every other organization in the area is making the same call. Knowing your actual runtime, and having a fuel contract that specifies priority rather than best effort, converts a number on a spec sheet into something you can rely on.

Resilience Planning Is Not Only About Cyber

Availability threats do not arrive exclusively through the network, and treating power as a facilities matter separates it from the resilience work where it belongs. The Cybersecurity and Infrastructure Security Agency addresses infrastructure resilience broadly, including the dependencies between sectors and the reality that power underlies essentially every other capability an organization has. Applying the same rigor to a power dependency that you would to a single-vendor dependency in your stack surfaces the same kinds of questions: what happens when it fails, how long can we operate without it, how do we know the mitigation works, and who is responsible for confirming that.

Grid Conditions Are Getting Less Predictable

The context has shifted in ways worth acknowledging. Extreme weather events affecting distribution networks have become more frequent across much of the country. Demand growth, driven substantially by data center construction and electrification, is placing pressure on grid capacity in specific regions. Public safety power shutoffs are now a normal operating practice in some fire-prone areas, meaning a planned multi-day outage is a scenario businesses in those regions should treat as likely rather than remote. Whatever view one takes of the causes, the planning implication is the same: outages of meaningful duration are a reasonable thing to design for.

Get the Load Calculation Right

The most common technical error is sizing based on nameplate ratings, which produces a generator considerably larger and more expensive than required, or worse, one sized for equipment that has since been replaced. Actual measured load matters, as does startup surge for motors and cooling equipment, and so does the decision about what genuinely needs to stay up. Very few organizations need every circuit backed. Deciding deliberately what is essential, what can run degraded, and what can simply stop is both an engineering exercise and a business one, and doing it properly usually reduces the cost of the solution.

Own It Before You Need It

The practical step is unglamorous: find out what backup power your facilities actually have, when it was last tested under load, how long it can run on stored fuel, and what specifically it covers. Most operations teams cannot answer those four questions about their own building. Getting the answers takes an afternoon, and it converts an assumption into either a documented capability or a known gap you can plan around. Either outcome is better than discovering the truth at two in the morning during a storm.