Hero image for "The Runbook You Wrote in June Didn't Mention the Day Four Providers Failed at Once"

The Runbook You Wrote in June Didn't Mention the Day Four Providers Failed at Once


On September 4, 2026, OpenAI, Anthropic, Google, and xAI all degraded within the same few-hour window, according to reporting compiled by baeseokjae's postmortem analysis. Anthropic's incident started at 9:23am ET and took about 15 minutes to diagnose, with impact resolved by 12:16pm ET. OpenAI's routing error hit ChatGPT and Codex at 7:43am PT, got a partial fix 30 minutes later, then suffered a second elevated-error period before full resolution around 12:55pm. xAI's Grok outage traced back to its Memphis compute center and ran roughly 3.5 hours across web, mobile, and API surfaces.

Here's the part that should bother anyone who built a "multi-model fallback" for resilience: it didn't matter. Four supposedly independent providers went down in the same window because they share cloud regions, network paths, and — increasingly — the same accelerator supply chain, per the same postmortem. If your on-call runbook's fallback step is "switch to the other provider," you need a plan for the day the other provider is also down.

A Runbook Entry Needs Four Parts, Not a Wall of Prose

Most teams write LLM runbooks the way they write database runbooks, and that's the first mistake. Classic latency and error-rate alerting doesn't catch the failures that actually hurt an LLM feature — quality decay, provider degradation, silent context growth — because none of those produce a slow request or a 500, according to FDE Notes. You can have green dashboards and a model that's quietly worse than it was yesterday.

The structure that actually survives a 3am page, per the same source, has four parts: the signal, the first action, the escalation condition, and the fallback. The first action has to be mechanical — not "investigate," but "check this specific page, compare this number to its value from a week ago, and if the delta crosses this threshold, do this." Fallbacks belong in the document because they're pre-made decisions, not improvisations: route to a different provider, drop to a smaller model, shorten the context, disable reranking, or kill the feature entirely. Each one needs a trigger and a named person authorized to pull it — not a Slack thread where three people debate whether now is the time.

That last point is where the September 4 event actually teaches something new. "Route to a different provider" is a fine fallback entry — until the trigger condition is "my provider is down" and the real condition should be "my provider's class of infrastructure is down." A runbook that doesn't distinguish those two will have an on-call engineer burning fifteen minutes trying three providers in sequence while all three fail for the same underlying reason.

Static Documents Rot; the Procedure Needs to Be a Service, Not a Memory Test

There's a sharper argument from ShiftMag: the standard prose runbook is institutional debt from the moment it's written. It's accurate for two weeks, then the dashboard links point at a tool you replaced and half the commands are deprecated. Nobody rehearses it, so the first real execution of the procedure is the incident itself. I'd go further for AI features specifically — an LLM runbook ages faster than a database one, because your fallback chain, your provider list, and your acceptable degraded modes change every time you touch the model config, not just every time you touch the infra.

The fix isn't a better wiki page. It's pushing the deterministic parts of the procedure into code that runs them, while keeping judgment calls with a human. Akash Talole's framework for runbook automation splits this into three tiers: fully automated for reversible, low-blast-radius actions; AI-assisted with mandatory approval for anything harder to reverse; and AI-informed, human-executes for everything else. The automation only works if the underlying runbook steps were already precise enough to execute without interpretation — which is the same four-part structure FDE Notes is describing, just compiled instead of read.

If you're going to let an agent touch production during an incident, keep it read-only by construction. CheatCoders' pattern allowlists kubectl get/describe/logs and blocks apply/delete/exec at the command-classifier level, so the agent can gather evidence and draft a remediation with cited sources, but a human has to approve and submit before anything executes. The approval step has to survive someone half-awake at 3am tapping a Slack button — "approve rollout restart on payments-api, change CHG-9182" is a decision a tired engineer can make correctly. A toggle labeled "auto-remediate" is not.

Test the Switch Before You Need It

None of this works if the fallback path is untested. FDE Notes puts a number on it: exercise the degraded modes and the feature kill switch during business hours, once a quarter, and record how long each takes. The failures that show up are usually boring — a permission only the deployment system has, a config path nobody exercised outside production. An hour a quarter is cheap insurance against finding that out during the next correlated outage, whenever it lands.