The SRE Book Era: Error Budgets Meet Cascading Failure

The SRE Book Era (Jul 2015 – Sep 2016) Filed September 20, 2016 — one year to the day after the DynamoDB cascade that defined this window. Google published Site Reliability Engineering in 2016 and gave the industry a shared vocabulary: SLOs, error budgets, toil, blameless postmortems. At the same time, the era’s biggest outages kept showing the same pattern. Systems broke not from the original fault, but from how they tried to recover. ...

September 20, 2016 · July 2015 – September 2016 · Retrospective

Shared Fate: Heartbleed, Mass Reboots, and the Limits of Cloud Trust

Shared Fate in the Cloud (Apr 2014 – Jun 2015) Filed July 12, 2015 — the weekend after the NYSE, United Airlines, and the WSJ all fell over on the same Wednesday. This is when everyone learned that moving to the cloud means sharing your provider’s mistakes too: their patches, their config pushes, their operators' typos. Security incidents also started getting the same careful postmortem treatment as plain downtime. The incidents that defined the period Heartbleed, April 2014. The OpenSSL bug that forced mass certificate rotation across the internet. The real operational lesson was that almost nobody had a list of where TLS actually terminated, so “patch and rotate” took weeks. Shellshock in September repeated the drill for bash. Joyent, May 2014. An operator running a routine update rebooted an entire data center of customer systems with one command. Joyent’s postmortem was direct: the problem wasn’t the operator, it was that the tooling let you target a whole datacenter with no confirmation. AWS Xen reboot, September 2014. AWS rebooted a large share of EC2 instances to patch a Xen bug before it was disclosed. Customers who had designed for instance failure barely noticed. Those who hadn’t found out the hard way which servers they’d been treating as pets. Microsoft Azure Storage, November 2014. A performance fix went out globally, skipping the usual staged rollout, and an infinite loop in the blob frontends took down storage across regions. Microsoft’s postmortem admitted the team had gone around its own rollout policy. It’s one of the most-cited config-change postmortems there is. NYSE, United Airlines, and the WSJ, July 8, 2015. Three unrelated same-day outages that everyone assumed were connected. A good example of how reliability failures turn into a news cycle. What the postmortems reveal Configuration changes became the leading cause. The Azure writeup captured a pattern that still dominates postmortems: the code was fine, and the rollout of a config flag was what broke things. “All deploys are staged, including config” started showing up in action items. ...

July 12, 2015 · April 2014 – June 2015 · Retrospective

The Blameless Awakening: How Postmortems Became Engineering Culture

The Blameless Awakening (Jan 2013 – Mar 2014) Filed April 3, 2014 — the week healthcare.gov closed its first open enrollment, ending the saga that taught the industry to read postmortems. In early 2013, few companies published real postmortems. Outages were treated as PR problems, not things worth explaining. That changed fast. By 2014, writing an honest postmortem was becoming normal, and this 15-month window is where the shift started. ...

April 3, 2014 · January 2013 – March 2014 · Retrospective