I said last time that the world’s ledger was resuming its ordinary cadence. The world heard me and put four cloud providers on the board in nine days.
The run, in order. July 16: AWS CloudFront couldn’t load updated network config into the fleet handling private VPC origins, and three and a half hours of 5xx errors landed on everyone downstream — Hugging Face, Canvas, Blackboard, a long tail of SaaS whose customers had no idea they were CloudFront customers. July 20: a power disturbance in a Google Cloud zone in the Netherlands defeated the backup power and the cooling, and the machines shut themselves down rather than cook. Fifteen hours. July 23: an Azure maintenance process in West US removed more network routes than it meant to — five hours where the workloads were fine and reaching them wasn’t. July 24: Oregon lost internet connectivity for eighty minutes and took Apple Pay, Reddit, Hulu, DoorDash and PlayStation Network with it. That last one was AWS’s third distinct incident in eleven weeks, which is the part the trade press led with, correctly.
The pattern is the boring one, again. Three of the four were the management layer, not the capacity layer: config that failed to distribute, maintenance that withdrew too much, automation doing precisely what it was told. Nobody ran out of servers. I have been typing some version of this sentence since a disk filled up on me in 2013, and the only thing that changes is the altitude.
What it cost us: nothing, and I want to be careful about how proud I’m allowed to be of that. The degraded-mode program went into its audit quarter last month, and it earned the audit — our critical path survived every one of those nine days because it’s built to survive any one of cloud, CDN, or model provider going away. That’s the funded outcome of the concentration-risk memo, and the memo was written by the outage autumn, not by me. I mostly formatted it.
The one that wasn’t software is the one I keep turning over. Fifteen hours, the longest single outage of the month, and the cause was electricity and heat. Every dependency review I’ve ever run stops at a vendor name. None of them ask where the power comes from, or what happens when the backup that’s supposed to cover the power also needs the cooling that the power was running. We’re adding the question. I expect the answers to be vague.
And the fortnight’s real news at home: Fantasy Premier League draft night, year fifteen. Null Pointer Exception enters the season with a squad assembled by a human being for the first time since the World Cup ate my attention span, and the group chat has ruled the Monte Carlo engine still banned under the constitution. Some things survive every migration.
TIL: CloudFront never went down. VPC Origins went down — one feature, and only the customers using it noticed, right up until the SaaS products they depended on went with it. Our dependency inventory lists vendors. It does not list which of each vendor’s features we’re actually standing on, which means it’s been half an inventory this whole time. Proverbs 353: know which parts of your provider you depend on, not just which provider.