OpenAI shipped o1-preview Thursday (September 12th), and the file marks it as the year’s genuine capability-architecture event (the watch-list item, “systems that think rather than chat,” arriving via a different door than expected): the model reasons before answering, chain-of-thought generated at inference time, hidden from the user, sometimes for tens of seconds, and the benchmark deltas are not incremental (83rd percentile on AIME math versus GPT-4o’s ~13th; PhD-level science questions crossing expert baselines; competition-code performance jumping a league). The structural insight the whole industry is now metabolizing: this is a second scaling axis, capability purchasable at inference time (more thinking tokens per question) rather than only at training time (more parameters per model), which re-prices everything downstream: the compute-scarcity trade extends from training clusters to serving fleets (thinking is expensive per-query now, margin structures and latency budgets both re-open), the eval discipline must handle non-deterministic depth (our golden sets now need difficulty tiers: when is a 30-second answer worth 30 seconds?), and the agent-spring’s failure taxonomy (goal drift, decomposition spirals) meets a model that does its own decomposition internally, with the reliability curve to be discovered in production, per tradition. The file’s calibrated note: the hidden chain-of-thought is also a transparency regression by design (the reasoning is the moat and the safety surface, and users see neither, the constants-file politics now includes the thoughts themselves), and the “reasoning model” framing will be both earned and oversold simultaneously (both-things doctrine, permanent resident).

The fortnight’s other altitude record, literal: Polaris Dawn, Jared Isaacman’s crew performed the first commercial spacewalk Thursday (same day as o1; the future ships in batches now), in SpaceX-designed EVA suits, from a Dragon at the highest crewed orbit since Apollo, the barnstorming-era file graduating another rung (the suit is the product: scalable EVA capability built outside a government program for the first time, and the dissimilar-redundancy ledger notes the asymmetry of the same fortnight’s Boeing news without cruelty, merely arithmetic). Apple’s iPhone 16 event (September 9th) shipped hardware built around Apple Intelligence features that ship… later, in stages (the system-layer bet now visibly gated on the eval problem at consumer scale, the demo-to-deploy gap is everyone’s gap this year), and the EPL opened its season with the group chat’s Fantasy PL draft concluding, for the twelfth consecutive year, in litigation (Null Pointer Exception’s roster: goalkeeper-diversified, forward-fragile, hope-heavy; the lineage unbroken).

TIL: inference-time compute scaling laws. Accuracy as a function of thinking-token budget, the emerging charts showing log-linear gains that mirror training-compute curves. The bitter lesson now has a per-query price tag, and the industry’s next margin war is literally over the cost of thought (the topology, extended into the request path).