Patch Notes #312 — The Day us-east-1 Forgot Its Own Name
Monday, October 20th: AWS us-east-1 went down for the better part of a day, a latent race condition in DynamoDB’s automated DNS management produced an empty DNS record for the service’s regional endpoint, the automation could not self-repair (the planner and enactor desynchronized; the fix required humans to disable the automation and restore state manually), and the cascade ran the full old syllabus at 2025 scale: DynamoDB’s resolution failure propagated into EC2 instance launches, network load balancers, Lambda, and the seventeen-service dependency web that us-east-1 has been since this archive’s first outage entry (the S3 typo, 2017, the file pulls the thread taut: eight years, the same region, the same lesson, the blast radius grown by an order of magnitude because the dependence grew while the topology didn’t diversify). Snapchat, Fortnite, Signal, banks, airlines, smart beds (the IoT ledger achieving its most absurd citation: mattresses with cloud dependencies stuck at heating settings), ~1,000+ companies filed impact, and the postmortem (published with AWS’s customary specificity) delivers the era’s central finding once more, now at its own source: the automation that manages the system is the system (the composed-failsafes doctrine, the GCP null-pointer, the channel file, the archive’s decade-long thesis now demonstrated by all three hyperscalers within eighteen months, each at the control plane, none at the capacity layer: the machines were fine; the management of the machines ate itself). Our own dependency-tiering held (the paper-runbook drawer was not needed but was checked, which is the drill’s entire point), and the fortnight’s industry-wide action item is this archive’s oldest sentence wearing its newest costume: know what you depend on, including what your automation depends on, including what it depends on when it’s wrong. ...