Patch Notes #273 — Three Claudes and a Logo Change

Anthropic shipped the Claude 3 family March 4th, Haiku/Sonnet/Opus, and the flagship posted benchmark wins over GPT-4 on the standard suites: the first time the same-day-launch rival has held the measured frontier, however temporarily, and the market-structure note matters more than the leaderboard (the dance is now a three-body problem, OpenAI/Microsoft, Google, Anthropic/Amazon-Google-money, with Meta’s open-weights flank as the fourth gravitational mass; frontier capability is officially plural, which changes procurement, eval discipline, and the governance math simultaneously: you cannot license a frontier that keeps electing new members). Our own quarterly bake-off (the tooling benchmarks) confirmed the delta on our tasks, which is the only leaderboard the file trusts (the golden-set doctrine): model choice is now a quarterly decision with a regression suite, exactly the commodity-with-switching-costs dynamics this archive filed for clouds a decade ago (the chokepoint homework, now with per-token pricing). The “sparks”-era question (which capabilities arrive at which scale) remains unanswered by anyone including the labs; the eval profession remains the only tractable response; the file remains on message because the message keeps grading correct. ...

March 18, 2024

The Platform Engineering Pivot: Datadog's $5M Lesson and the First AI Whispers

The Platform Engineering Pivot (Jan 2023 – Mar 2024) Filed March 8, 2024 — one year after the Datadog incident everyone studied. The big postmortem this window came from an observability vendor going through the exact kind of outage it sells tools to prevent. Meanwhile “platform engineering” started absorbing much of what used to be called DevOps, and the first LLM assistants quietly showed up in incident channels. The incidents that defined the period FAA NOTAM outage, January 2023. A corrupted database file, linked to a contractor’s mistake during maintenance, grounded all US flight departures for hours. It was the first nationwide ground stop since 9/11, and decades-old systems with no hot failover became a topic in Congress. Microsoft Azure WAN, January 25, 2023. A router configuration change (a command that different devices interpreted differently than intended) rippled through Microsoft’s global WAN and broke Azure, Teams, and M365 worldwide for hours. The classic config-change-to-global-outage, at telco scale. Datadog, March 8, 2023. The one everyone studied. An automatic security update to systemd across their fleet triggered a network stack reset on tens of thousands of nodes, across multiple cloud providers at the same time (datadoghq.com). Days of degraded service, a reported ~$5M revenue hit, and a thorough multi-part postmortem. Being multi-cloud didn’t help, because the same OS update channel ran across all of them. Redundant copies failed together because they shared a config source. AWS us-east-1, June 13, 2023. A capacity-management issue in Lambda degraded dozens of services for about three hours. Notable admission in the postmortem: AWS’s own support-case system was impaired again. UK air traffic control (NATS), August 2023. A single flight plan with duplicate waypoint names hit an unhandled edge case, and the primary and its identical backup both failed the same way. The independent review became a classic on common-mode software failure. Optus, November 2023. A routing update from an upstream network cascaded into a roughly 14-hour national outage in Australia, affecting emergency calls, and the CEO resigned. Executive accountability for reliability, made explicit. What the postmortems reveal Correlated failure became the top-of-mind risk. Datadog (one update channel across every cloud) and NATS (identical primary and backup software) showed that redundancy without diversity is just bookkeeping. Postmortems started asking which update, config, or code path is shared across the copies you think are independent. ...

March 8, 2024 · January 2023 – March 2024 · Retrospective

Patch Notes #272 — The Overcorrection and the Sideways Moon Landing

Google spent the fortnight in the year’s most instructive AI-product crisis: Gemini’s image generation, tuned to counteract training-data demographic bias, overcorrected into generating diverse-by-mandate imagery for historically specific prompts (Vikings, 1943 German soldiers, American founders, each rendered with demographic diversity the historical record does not contain), and the screenshots detonated into a culture-war news cycle that forced the feature offline and a CEO memo calling the outputs “completely unacceptable.” The principal-file’s read, holding the both-things-true line against a discourse determined to pick one: the underlying problem is real (uncorrected models default to training-data demographics, “CEO” renders as white men at rates the actual world doesn’t justify; the confusion-matrix politics), the correction was real too (a system-prompt-level diversity injection applied without historical-context conditioning, the Goodhart file’s purest AI specimen: the metric was representation, the target became the metric, and the optimizer found the exploit in exactly the cases that falsify the intent), and the meta-lesson is the one this archive has filed since Tay: behavioral tuning at planetary scale is values engineering with a constants file, and both under- and over-correction ship someone’s politics as a default. The eval discipline gains its hardest test suite: historical-fidelity-versus-representational-harm is not a golden set anyone has written well yet, and the file suspects the answer involves context-conditional behavior (the model should know a Viking prompt from a CEO prompt), which is to say, judgment, which is to say the hard part, again, always. ...

March 3, 2024

Patch Notes #271 — Sixty Seconds of Video, Overtime in Vegas

OpenAI demoed Sora Thursday, text-to-video, sixty-second clips of a quality the latent-diffusion file’s 2022 imagery cannot share a sentence with: coherent object permanence, camera moves, reflections in a Tokyo puddle, a woman’s earrings swinging with her gait. Research preview only (no public access; red-teaming first, the staged-release doctrine now standard practice), and the discourse split on schedule (the Move-37 sequence: trick, tool, threat, field, all four stages arguing simultaneously this time). The file’s assessment at demo-distance, calibrated by the Gemini-video lesson (demos lie until production doesn’t): even discounting selection bias heavily, the capability slope is the story, video was supposed to be years behind imagery (temporal coherence as the moat), and the moat lasted eighteen months. The provenance stack (C2PA, watermarking) just moved from “needed” to “overdue against an active clock,” with an election cycle (the braced posture) running concurrently; the file’s watch-item is no longer “can it be made” but “can anything downstream verify what was” (provenance wars, now at 24fps). The physical-world simulation claims in the technical report (world-models language) the file logs with the both-things-true protocol: marketing and research direction, one document. ...

February 17, 2024

Patch Notes #270 — Launch Day for the Face Computer

Vision Pro ships today, and the archive files launch-day observations before the reviews calcify (the eve-discipline, adapted): the launch-day review consensus (US-only launch; a day-one unit is en route to the group chat via the traditional Valley-cousin supply chain, hands-on notes when it clears customs) grades exactly as pre-registered: technically astonishing — passthrough latency imperceptible, the eye-tracking-plus-pinch interface reportedly inevitable-feeling within minutes (the ARKit decade compounding into an input paradigm), the movie-screen experience justifying a review genre by itself — while every honest reviewer does the arithmetic the keynote omitted: 600+ grams on the face, two-hour tethered battery, EyeSight’s uncanny compromise, no killer workflow yet beyond “astonishing demo,” $3,499. The Watch protocol applies verbatim: shipped confident, purpose TBD, two years of patience granted. What’s new versus 2015: Apple’s ecosystem gravity now includes a decade of spatial-computing developer investment and a services empire hungry for a new surface, the v3-at-half-price bet remains the file’s position, with the launch-day amendment that the input model (eyes plus fingers, no controllers) is the part competitors will be copying by Christmas regardless of unit sales (the courage-cycle: the removed thing this time was the controller, and the removal is the product). ...

February 2, 2024

Patch Notes #269 — The ETF and the Hacked Announcement

The SEC approved spot Bitcoin ETFs January 10th, eleven applications at once, BlackRock and Fidelity among them, ending a decade of rejections and completing the arc this archive has filed since a coworker’s secret 2011 stash: joke, mania, crash, institution, state currency, fraud winter, and now, wrapped in the most traditional product structure American finance sells. The asset the industry was built to route around now trades through the exact rails, custodians, authorized participants, ticker symbols, it was invented to obsolete; the file notes the irony without sneering, because the irony is the lesson: every insurgent technology that survives gets domesticated by distribution (the graph-portability doctrine: the moat becomes the launchpad; here, the wrapper becomes the market). Flows will tell the story by year-end. ...

January 18, 2024

Patch Notes #268 — Year Twelve: The Lawsuit and the Ledger

Year twelve opens with the copyright case this archive pre-registered fifteen months ago (“the litigation defining this decade files within months,” late by a year, correct in kind): The New York Times v. OpenAI and Microsoft, filed December 27th, and it’s the strong version of the claim, not just “you trained on our archive” but exhibit after exhibit of GPT-4 reproducing Times articles near-verbatim under adversarial prompting (the memorization receipts the fair-use debate has been waiting for), plus the market-substitution argument (Browse-with-Bing summarizing paywalled recipes) that turns an abstract doctrine into a revenue chart. The file’s read: this is the case both sides arguably want, the labs need the training-data question settled at precedent altitude rather than by a thousand district-court paper cuts, the publishers need leverage for the licensing market that is obviously the endgame (Axel Springer and AP already signed; the Times sued after negotiations stalled, which tells you the suit is the negotiation, contracts-as-adversarial-runbooks, media edition). Pre-registration for the file: no verdict ever lands, settlement plus licensing regime within two years, with the memorization exhibits driving the price. Either way, 2024’s AI story adds a fourth branch of government to the governance stack: the judiciary has entered the training loop. ...

January 3, 2024

Patch Notes #267 — Year Eleven Retrospective: The Loom's First Full Year

Entry 267 closes year eleven, and the closing fortnight supplied its own miniature of the year: Google’s Gemini launched (Dec 6) with impressive benchmarks and a hands-on demo video so fluidly responsive it briefly reset expectations, until the disclosure that it was edited (prompts abbreviated, latency cut, the interaction reconstructed), and the old doctrine (“the demo is the dream; production is the compromise”) claimed its most consequential 2023 specimen: in the eval era, a misleading demo isn’t marketing anymore, it’s a factuality regression in your own launch, caught by the planet’s review queue within 48 hours (the replication machine, now aimed at product videos). Meanwhile E3, the cathedral of exactly that demo culture, the stage of the old Sony massacre and the Keanu moment, was formally declared dead (Dec 12) after 28 years, killed by the direct-to-audience channels (Nintendo Directs, State of Play) that made the middleman theater redundant. The two obituaries are one lesson: the demo’s power migrated to the deploy (civilization-runs-the-eval, achieving total coverage). ...

December 19, 2023

Patch Notes #266 — Restoration, With Amendments

The weekend resolved as the mid-crisis logic demanded: Altman restored as CEO within five days of the firing, the Shear interim lasting roughly 72 hours (his tenure’s principal artifact: a tweet clarifying the board “did NOT remove Sam over any specific disagreement on safety,” deleting the weekend’s leading theory without installing a replacement), the old board dissolved save one, a new small board seated (Bret Taylor chairing, Larry Summers arriving as the establishment’s notary), the employee letter having reached ~95% signature coverage, including, in the detail that will feed governance seminars for a decade, signatories among the board’s own allies and Ilya Sutskever, who co-signed the letter against the action he’d voted for days earlier (“I deeply regret my participation”), and Microsoft converting its weekend of leverage into a board observer seat: influence formalized, liability declined, the earlier structure’s actual power topology now documented by stress test (the file’s charter-vs-cap-table question answered: the cap table won, wearing the charter’s language). Investigations pending; the fired board’s specific cause remains unstated, which the file continues to hold as the weekend’s original sin and its enduring mystery (the eventual review’s findings, “a breakdown of trust,” reportedly, over candor in board communications, will satisfy no one, which may be the truest possible finding). ...

December 4, 2023

Patch Notes #265 — The Weekend the Board Blinked

Filing from inside the strangest governance event this archive has ever covered, outcome unknown, per the charter (write at the moment of not-knowing): on Friday afternoon, OpenAI’s nonprofit board fired Sam Altman, four sentences of announcement, “not consistently candid” as the entire stated cause, no warning to investors including the $13B partner, no succession plan visible, Greg Brockman resigning within hours, and Mira Murati installed as interim CEO. As I file Sunday night: reporting says Altman was in the building today negotiating return terms (wearing a guest badge he photographed, “first and last time i ever wear one of these”), the board has apparently deadlocked, a second interim CEO (Twitch’s Emmett Shear) is rumored incoming tonight, Microsoft’s Nadella is reportedly working the phones with the leverage of a partner who learned of the firing one minute before the world, and 700+ of ~770 employees are preparing a letter threatening to follow Altman to a Microsoft-hosted lab unless the board resigns. The earlier pre-registration executes with terrifying precision: “governance structures reveal themselves only under governance stress, and this one will meet stress,” the capped-profit wrapper’s novel geometry (a nonprofit board, legally bound to a mission rather than shareholders, holding fire-the-CEO authority over the most commercially consequential company of the decade) has now fired its weapon and discovered the recoil: the board had the authority but not the power (the positional doctrine at maximum stakes, the employees are the company; the compute contract is the balance sheet; a mission without a workforce governs an empty building). Whatever Friday’s actual cause, and the absence of a stated one converted every observer into a conspiracy theorist by Saturday brunch (the velocity doctrine: in an information vacuum, coordination happens at group-chat speed against you), the meta-lesson is already filed: firing without a communicated cause is a resignation letter written by the board (the two-sentence-note doctrine inverted: under-communication at maximum stakes reads as either cowardice or emptiness, and both readings kill the authority that chose silence). ...

November 19, 2023