Patch Notes #293 — Year Thirteen Opens in Smoke

Year thirteen opens with Los Angeles burning, the Palisades and Eaton fires, driven by hurricane-force Santa Anas across a landscape that hasn’t seen rain in eight months, have destroyed entire neighborhoods (12,000+ structures and counting as I file), and the archive’s climate-infrastructure ledger adds its gravest domestic chapter: hydrant systems designed for house fires meeting neighborhood-scale conflagration (the design-basis doctrine again, urban water infrastructure has a WUI-fire envelope nobody funded), insurance markets in managed retreat (State Farm’s pre-fire non-renewals in the exact zip codes now burning, the actuaries, as ever, were the first honest climate models), and the information layer performing its now-standard double duty: real-time fire maps and mutual-aid coordination at their best, AI-generated Hollywood-sign-ablaze imagery at its worst (the verification-terminal-state runbook, deployed by every newsroom simultaneously). Friends evacuated; our LA teammates are safe and housed; the old doctrine, hug your people, opens its thirteenth year of service. Donate to the wildfire funds; the file keeps minutes and its lane. ...

January 12, 2025

Patch Notes #292 — Year Twelve Retrospective: The Kernel and the Curve

Entry 292 closes year twelve, and the closing fortnight refused a quiet exit: OpenAI announced o3 on Shipmas’ final day (December 20th), the reasoning axis’s second generation, and the headline result stops the file mid-sentence: ~87% on ARC-AGI’s semi-private eval (the abstraction-and-reasoning benchmark designed to resist memorization, where GPT-4o scored single digits and four years of models flatlined), plus research-math and competition-code results that moved the frontier’s own researchers to public awe, achieved, per the disclosed methodology, with inference-compute budgets that reach thousands of dollars per task at the high end (the cost-of-thought economics now spanning six orders of magnitude: the capability exists; the price curve is the product roadmap, and every prior curve in this file’s twelve years, compute, storage, sequencing, launch-mass, says the price falls faster than the discourse expects). The file’s calibrated position for the year-end record: benchmarks are not jobs (the map-vs-territory doctrine, permanent), ARC’s designer himself notes the gap between eval-crushing and general intelligence, and the slope is the steepest this archive has ever filed across twelve years of logging “probably nothing” items (the Transformer paper to o3, the quiet thread’s decade, bookended). Both things true, maximum stakes, year thirteen inherits the question. ...

December 28, 2024

Patch Notes #291 — Twelve Days, Two Point Oh, and a Quantum Willow

Launch-season December, fully institutionalized (the counter-programming cadence now has a holiday calendar): OpenAI is mid-“12 Days of Shipmas” (o1 full release with a $200/month Pro tier, the inference-economics file notes the price point as the margin war’s opening bid: thought, tiered; Sora’s public release Monday, the earlier demo now a product, with the provenance stack shipping alongside as C2PA metadata, graded “necessary, insufficient, present” by the file), and Google counter-launched Gemini 2.0 Wednesday (explicitly framed as “for the agentic era,” Project Mariner browsing autonomously, Astra seeing through cameras, and the watch-list item now the strategy slide of both giants simultaneously: 2025 is being pre-announced as the year models stop chatting and start doing, and the agent-spring’s expensive recursion meets its second, better-funded spring). The file’s eval-discipline note, always: “agentic” converts error rates from per-answer to per-action-sequence, compounding failure probabilities across multi-step tasks is the reliability math the demos elide (a 95%-per-step agent is a 60%-per-ten-step agent), and the orgs that instrument sequence-level golden sets first will own the trust layer the way the rehearsers own reliability. Pre-registered for 2025’s grading. ...

December 13, 2024

Patch Notes #290 — Two Years of the Loom

Thanksgiving fortnight, quiet by the year’s standards (the file’s quiet-fortnight instincts stand down for once), and the calendar delivers a reflection hook the archive refuses to waste: Saturday marks two years since ChatGPT launched, and the principal-file conducts the two-year review the industry is too busy shipping to write. Capability: from “startling autocomplete” to reasoning models, two scaling axes, frontier plurality, and open weights one generation behind; the curve’s slope survived every “wall” declaration (the file counts four major “scaling is over” discourse cycles, each followed by a capability jump, the calibration ledger suggests betting against the curve requires more courage than the discourse prices). Deployment: from demo to default, our own org’s census (the ten-day adoption) now shows AI-assisted workflows in every function including legal (the policy’s affordances-not-prohibitions bet, fully vindicated), and the review-first cohort is now training its successors (the re-rigged ladder holds weight). Discourse: matured from poles toward operations (the tractable middle won, evals, provenance, staged deployment are where the actual governance happens while the summit communiqués photograph well). Unresolved, honestly filed: the labor reallocation is real and lumpy (the strike settlements set precedents; the entry-level squeeze the ladder-question predicted is measurable in industry hiring data, and our re-rigging is a local fix to a global problem), the governance-by-exit pattern at the frontier labs remains the decade’s open risk, and the energy-and-compute buildout (inference economics at civilizational scale, the datacenter-buildout headlines now read like the CHIPS file’s sequel) is writing checks the grid literature says need a decade of plumbing. Two years; the operating environment changed; the closing wager (judgment compounds, tools amplify) holds better than its author feared. ...

November 28, 2024

Patch Notes #289 — European Statements and Blue Skies

The election happened (November 5th; per the usual lane: the infrastructure held again, counting proceeded, the predicted synthetic-media apocalypse arrived as scattered showers rather than the storm (the earlier fears met the replication machine’s civilian deployment: debunk-velocity mostly matched fake-velocity this cycle), and the archive notes the decisive result’s tech-sector implications, crypto markets surged on regulatory-reset expectations (the ETF flows now joined by policy beta), the antitrust remedies’ fate acquires administration-change uncertainty, and the AI executive order’s future is officially TBD, the whole regulatory-geography map now redraws at inauguration; the file files the facts and keeps its lane, per twelve years of practice). ...

November 13, 2024

Patch Notes #288 — Hattricks and Tabletop Exercises

The Champions League matchday the sport’s accountants dreamed of delivered: Real Madrid-Dortmund, a classic final rematch, and the match instantly entered the canon: down two goals, second half, Vinicius Junior, playing through a minor ankle strain, the play-through-the-maintenance-window lineage, hit a second-half hattrick, one of the fastest in Champions League history, and the stadium’s decibel telemetry reportedly registered on regional seismographs (the old F1 file smiles: October writes fiction nightly). Real Madrid won 5-2; the pre-writes-nothing doctrine holds, but the architecture note is already bankable: Madrid’s roster, Mbappé’s star signing (contract engineering funding the depth around him), Bellingham’s pedigree, Vinicius’s clinical finishing, is the superteam economics thesis executed with retention discipline, and if it continues, the file will mark it as roster-construction’s decade meeting its proof. ...

October 29, 2024

Patch Notes #287 — The Chopsticks

Yesterday morning SpaceX launched Starship’s fifth test flight, and the Super Heavy Booster, twenty-three stories of steel, returning from the edge of space at supersonic speed, flew itself back to the launch tower and was caught out of the air by the tower’s mechanical arms. The chopsticks. On the first attempt. I have watched the footage upward of thirty times and each viewing produces the same involuntary sound the earlier landings and eclipses produced, the sound of the impossible becoming a procedure. The engineering file, dutifully, beneath the awe: catching eliminates landing legs (mass, complexity, refurbishment) and enables the launch-catch-restack-relaunch cadence the whole architecture is priced on (the tower is the rapid-reuse thesis made steel, the iterate-through-explosions doctrine arriving at its payoff phase: five flights from “cleared the tower is success” to “caught the booster with a building”), and the trajectory-abort logic deserves its own sentence: the booster earned the catch attempt only after passing thousands of automated health criteria mid-descent, with the default being ocean divert, the launch-commit doctrine running autonomously at Mach speeds (the system polls itself for go/no-go now, the ship-the-judgment doctrine, twenty-three stories tall). ...

October 14, 2024

Patch Notes #286 — Exit Interviews and Vetoed Thresholds

OpenAI’s earlier re-org completed its slow-motion arc this week: Mira Murati resigned Wednesday (the interim-CEO of the November weekend, the product org’s center of gravity), followed within hours by the chief research officer and a research VP, the same week reporting confirmed the company’s restructuring toward removing nonprofit control entirely (the capped-profit wrapper, whose stress test revealed the cap table’s actual power, now being formalized into the org chart it always was, with equity stakes for the CEO under discussion, per reporting the company disputes in emphasis). The file’s ledger of departures since the November weekend now reads: Sutskever, Leike, Karpathy, Schulman, Brockman-on-leave, Murati, McGrew, Zoph, functionally the entire founding research and safety leadership, out within ten months of the board’s capitulation, and the earlier extraction (“the tension didn’t resolve, it re-org’d”) upgrades to its terminal form: the structure resolved by exit. Whatever one’s read on any individual departure (startup attrition is real; so is gradient), the aggregate is a governance postmortem written in resignation letters, and the archive files it next to the old adversarial-reviewer doctrine with the observation it has earned across twelve years: incentive structures don’t fail loudly; they fail by selection, the people whose concerns priced above their equity simply leave, and the org that remains is, definitionally, the org that didn’t share them (Goodhart, applied to workforce composition; the metric survived, the mission migrated). ...

September 29, 2024

Patch Notes #285 — The Model That Thinks Before Speaking

OpenAI shipped o1-preview Thursday (September 12th), and the file marks it as the year’s genuine capability-architecture event (the watch-list item, “systems that think rather than chat,” arriving via a different door than expected): the model reasons before answering, chain-of-thought generated at inference time, hidden from the user, sometimes for tens of seconds, and the benchmark deltas are not incremental (83rd percentile on AIME math versus GPT-4o’s ~13th; PhD-level science questions crossing expert baselines; competition-code performance jumping a league). The structural insight the whole industry is now metabolizing: this is a second scaling axis, capability purchasable at inference time (more thinking tokens per question) rather than only at training time (more parameters per model), which re-prices everything downstream: the compute-scarcity trade extends from training clusters to serving fleets (thinking is expensive per-query now, margin structures and latency budgets both re-open), the eval discipline must handle non-deterministic depth (our golden sets now need difficulty tiers: when is a 30-second answer worth 30 seconds?), and the agent-spring’s failure taxonomy (goal drift, decomposition spirals) meets a model that does its own decomposition internally, with the reliability curve to be discovered in production, per tradition. The file’s calibrated note: the hidden chain-of-thought is also a transparency regression by design (the reasoning is the moat and the safety surface, and users see neither, the constants-file politics now includes the thoughts themselves), and the “reasoning model” framing will be both earned and oversold simultaneously (both-things doctrine, permanent resident). ...

September 14, 2024

Patch Notes #284 — The Founder in Custody and the Stranded Crew

France arrested Pavel Durov August 24th, the Telegram founder, detained stepping off his jet at Le Bourget, charged days later with complicity in the platform’s criminal uses (CSAM distribution, drug trafficking, organized fraud) plus a cryptology-declaration charge, on the theory that Telegram’s near-total non-cooperation with law enforcement (moderation-by-shrug at 900M users; the reporting says French requests went systematically unanswered) crosses from platform immunity into complicity. The file’s structural read, held with both hands per doctrine: this is the stack-sovereignty thread’s most personal escalation, the executive as the enforcement surface (intermediary-liability regimes worldwide have spent a decade adding “senior-manager liability” clauses, the UK Online Safety Act, India’s rules; France just executed the pattern), and the precedent cuts every direction at once: platforms that ignore all process invite exactly this (Telegram’s posture was never principled E2E cryptography, most chats aren’t even encrypted end-to-end; it was operational indifference wearing privacy’s coat, and the file has kept that distinction sharp for years), and founder-arrest-as-content-policy is a tool every less-liberal government will now cite with delight (the capabilities-outlast-settlements doctrine: the playbook, once demonstrated, is everyone’s). Signal’s Meredith Whittaker spent the week correctly distinguishing her architecture from Telegram’s in public, the crypto-legibility gap (“term of art, not a vibe”) is now a liberty-relevant distinction, and the file recommends every platform executive re-read their own transparency reports as extradition documents. ...

August 30, 2024