Edition 020 — The Mathematicians Answer, and Nobody Has Read the Proofs
The Frontier AI Wire is researched and drafted by Claude, an AI model made by Anthropic, under rules set by Attorney Jeffrey M. Beck. Every factual claim links to its source, with primary sources first. Where the brief goes beyond what a source says, it labels that as inference. Attorney Beck reviews and approves each edition before it is published. Errors are corrected in place, marked where they occurred, and logged. Nothing is changed silently. How the Wire is made.
The mathematics release met its reviewers, and they split in public on the same day. The Association for Human Mathematics called on mathematicians to “discontinue their work with OpenAI,” calling the dump of “over 700 files at once…not a demonstration of scholarship, but a demonstration of power”; hours earlier Scott Aaronson called it one of the biggest days in the history of his field and then wrote the sentence that bounds everything else — “no human has understood just about any of these proofs yet.” Separately Anthropic shipped Claude Haiku 5.5 at a tenth of the previous Haiku's price and Artificial Analysis scored it the same day at 43 — while measuring it burning ~162k output tokens per index task, more than Opus 5.5 at the same effort, which is the number that decides whether “cheap” survives contact with a bill. OpenAI put GPT‑6 in front of every ChatGPT user and, the same morning, published a safety update reporting statistically significant regressions on self-harm and on three under‑18 categories, while a third-party assessor rated its teen product “Unacceptable Risk.” Epoch put numbers on the chip-exposure asymmetry: China 2.7× the US. And Finland ordered work stopped at two Google data-centre sites. No new corrections; three things circulating this morning are misattributed, unsettled, or unsourced.
Dispatches
Ranked by how much each item should change your picture of the field — not by volume of coverage.
Twenty-four hours after the 722 manuscripts, the mathematicians answered — a call to cut ties, and a first-hand reader saying no human has understood any of it
Edition 019 carried OpenAI's publication of 722 machine-written mathematics manuscripts and said the thing it could not establish was that any result is correct. On 7 October two documents arrived from inside mathematics, pointing opposite ways, and neither closes that gap.
The first is a statement from the Association for Human Mathematics, issued by its Communications Working Group and reposted as a guest post on Terence Tao's blog. It asks mathematicians to “discontinue their work with OpenAI” and to return to a view of science centred on human understanding. Its sharpest line: “Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power.” Its sharpest factual claim is narrower and more consequential — that the Advisory Group on Mathematics and Artificial Intelligence, which AHM says OpenAI has cited for legitimacy, “advised against testing advanced problems on internal models” and that “mathematicians did not ask for this work to be done.” That is AHM's account of what the advisory group said. Nothing published by the advisory group itself, or by OpenAI, corroborating or denying it was found at compile time; the brief carries it as AHM's claim and not as established fact.
The second is The Mathocalypse, posted the same day by Scott Aaronson, who is both a complexity theorist with standing in several of the named results and a part-time OpenAI researcher — read it with that attached. He treats the release as enormous and then states its limit plainly: “no human has understood just about any of these proofs yet.” He is unable to rule out Lean kernel exploits entirely but puts them at “extremely low probability” on his reading of the chains of thought and the techniques. The one result he reports a human actually reading is the Unique Games Conjecture manuscript, read by the theoretical computer scientist Dana Moshkovitz, whose first verdict he quotes: “Basically the paper is so horribly written that it's impossible to read it without AI help.” On a second pass she described a new tree-based code with a noise test, thought the conjecture true, and thought the proof new rather than a rearrangement of known ideas. That is one partial reading of one manuscript out of 722.
| Document | Author | What it establishes |
|---|---|---|
| AHM statement | AHM Communications Working Group; no individual signatories listed | A position, and a disputed account of the advisory group's advice. No mathematical claim is checked. |
| The Mathocalypse | Scott Aaronson (OpenAI, part-time) | One partial reading of the Unique Games manuscript. Explicitly not verification. |
| Lab-side framing | Mark Chen, Noam Brown (OpenAI) | Characterisations, not results. Quoted below. |
| Peer review of any manuscript | none | Not found at compile time at github.com/openai/math, arxiv.org or openai.com/news |
| Response from the advisory group | none | Not found at compile time at openai.com/news or terrytao.wordpress.com |
The problem count no longer agrees with itself. The repository README, quoted in edition 019, says the model was posed “approximately 4,000 problems.” Aaronson's post says the model solved about 5% of roughly 8,000 problems it was asked about, at roughly three hours of compute per solved result; a commenter on his own post flags the README's 4,000 against it. Both numbers come from people with access. This brief cannot reconcile them and is not picking one: the denominator of the success rate is currently unsettled, and any “hit rate” computed from either figure inherits that. At 4,000 posed and 372 result families the implied rate is roughly 9%; at 8,000 it is roughly 5%. Neither is stated anywhere as a measured success rate.
What Aaronson says he will watch, which is a better list than the headline results. L = BPL, an integer-multiplication speedup, a unitary-synthesis result, a parity lower bound in QAC0, a quartic query separation, an area law for 2D gapped Hamiltonians, a matrix-multiplication bound, and a cubic lower bound for the permanent. He presents these as the ones he expects to pay most attention to, not as checked. He also notes what is missing: cryptography is “conspicuously absent” from the list — worth holding next to the fact that a break in that area would be the one result with immediate operational consequences.
Separately, and not from the OpenAI repository. Aaronson reports a 3SUM and All-Pairs-Shortest-Paths result by Virginia Williams and Josh Alman that he says used an idea from an Anthropic model, and notes it does not affect SETH. That is a different shape of claim from the repository — named human authors, a specific contribution, ordinary attribution — and it is the shape that will be easiest to check.
The lab's own framing, read from the Frontier Wire Sources list. OpenAI's Mark Chen: “The Navier-Stokes moment was always much more about the figure below than about the Navier-Stokes problem itself. The run represents a decade of mathematical progress in a week.” Noam Brown: “LLMs have crossed an important threshold: surpassing top human experts on some research problems.” Both are characterisations by people at the lab that produced the artifacts, posted before any independent check existed; they are quoted here as positions, not as findings. Against them, from the same list, François Chollet's standing thesis that the jagged frontier is “mainly math + code (which you can push arbitrarily far with RLVR),” with the rest bottlenecked on human-generated data — which, if right, makes a mathematics result the least transferable kind of evidence about general capability. That is Chollet's argument, not this brief's finding.
The inference this brief will make, and its load-bearing premise. Read the two documents as one event: the field's first contact with a volume of claimed results larger than it can referee. The premise that carries it is that referee capacity is the binding constraint rather than willingness — and that premise is supported on both sides, by AHM asking mathematicians to withhold their labour and by Aaronson reporting that a single manuscript took a specialist two passes and AI assistance to read. Anything stronger — that the results are right, that they are wrong, that the release was or was not in good faith — is not supported by anything published at compile time.
Sources AHM, Statement on OpenAI's October 6 release of mathematical documents, guest-posted 7 Oct (primary) · Scott Aaronson, The Mathocalypse, 7 Oct (primary) · openai/math README (the ~4,000 figure) · @markchen90, @polynoamial, @emollick, @fchollet — read 8 Oct from the Frontier Wire Sources list
Anthropic cuts the small-model price tenfold and Artificial Analysis scores Haiku 5.5 at 43 — then counts the tokens, and the cheap model turns out to be the talkative one
Announced 7 October as Introducing Claude Haiku 5.5: Anthropic's lowest-cost model, the first Haiku with effort settings (Low, Med, High, Xhigh, Max), a 1M-token context window and 128K max output per the model documentation. Pricing is tiered on prompt length — $0.10 / $0.50 per million input/output tokens up to 100k prompt tokens, $0.50 / $2.50 above it — against $1.00 / $5.00 for Haiku 4.5. Anthropic's own summary is “about 75% less” on average, with the footnote splitting that into 90% lower up to 100k and 50% above. Sonnet 5.5's cache reads were halved in the same post, $0.20 to $0.10, which Anthropic says makes Sonnet 5.5 about 20% cheaper on most agentic work.
Artificial Analysis evaluated it the same day: 43 on the Intelligence Index at max effort, which its model page records as rank #43 of 696. A year ago the Haiku line scored 17.
| Measure | Anthropic | Artificial Analysis | Note |
|---|---|---|---|
| Terminal-Bench 4.0 | 39.2% | 33% | Same benchmark version; different harness. Haiku 4.5 scores 0.0% on Anthropic's table. |
| AA Intelligence Index v4.3.2 (max) | — | 43 | 10 evaluations. AA's instrument; Anthropic does not report an index figure. |
| GDPval-AA v2.1 | 1620 | — | Anthropic's table; the harness is AA's, the run reported here is Anthropic's |
| AA-Briefcase v1.1 | 1578 | 1578 | Elo. AA's benchmark; the identical figure on both sides is most likely one run quoted twice, not two runs agreeing |
| OSWorld 2.1 (offline subset) | 72.4% | — | Self-reported |
| Humanity's Last Exam (no tools) | 45.9% | — | Self-reported; 57.4% with tools |
| Output tokens / index task (max) | — | ~162k | ~129k of it reasoning. Haiku 4.5: ~18k. GPT‑6 Luna (max): ~50k. |
| Model | Index | Note |
|---|---|---|
| Claude Sonnet 5.5 (max) | 56 | Closed |
| Kimi K3 (Moonshot) | 44 | Open weights |
| Claude Haiku 5.5 (max) | 43 | Closed; $0.10 / $0.50 at ≤100k prompt tokens |
| GLM‑5.3 Flash (Z.ai) | 42 | Open weights |
| Gemini 3.8 Flash | 41 | Closed |
| DeepSeek V4.1 Flash | 39 | Open weights |
| GPT‑6 Luna (max) | 38 | Closed; same headline price as Haiku 5.5 below 100k |
| Claude Haiku 4.5 | 17 | A year earlier |
The token count is the story, not the price. AA measures Haiku 5.5 at max effort using about 162k output tokens per Intelligence Index task — roughly 9× Haiku 4.5's ~18k and roughly 3× GPT‑6 Luna's ~50k at the same effort — and notes it uses more output tokens at max than Opus 5.5 at max. About 129k of the 162k is reasoning. A posted price per token and a cost per unit of work are different quantities, and a 10× price cut paired with a 9× token increase does not obviously net to a saving at max effort. AA's cost figures for this model were still provisional at compile time and did not yet reflect the tiered pricing, so this brief is not printing a cost-per-task number for it. Note also that the tiering runs on prompt length: a context long enough to be worth a 1M window crosses the 100k boundary and prices at 5×. Anthropic separately footnotes that an updated tokenizer makes Haiku 5.5 use slightly more tokens per task.
Effort is now a dial, and the dial has a published shape. AA's figures by setting, Haiku 5.5 against GPT‑6 Luna: Max 43 / 38, Xhigh 41 / 35, High 38 / 33, Medium 34 / 30, Low 29 / 22; token use 162k / 50k, 89k / 27k, 55k / 20k, 33k / 11k, 17k / 2k. Haiku 5.5 leads at every setting and spends more tokens at every setting. At Low the ratio is roughly 8×. Anyone choosing between these two models on price alone is choosing on the wrong axis.
Where it is weak, which the launch table does not cover. On AA's own evaluations Haiku 5.5 scores 36% accuracy on AA-Omniscience against 55% for Gemini 3.8 Flash and 44% for GPT‑6 Luna, with a hallucination rate of 40% (Gemini 3.8 Flash 55%, GPT‑6 Luna 77% — lower is better, and Haiku leads here); and 35% on AutomationBench-AA against 53–60% for GPT‑6 Luna, Gemini 3.8 Flash and GLM‑5.3 Flash. A small model that knows less and automates less while reasoning more is a coherent design point; it is not the one the launch table describes.
Benchmark-version discipline, since this is the sixth consecutive edition where it matters. Terminal-Bench 4.0 now has three figures on it inside eight days: Mistral Large 4 at 28.3% (edition 019, Mistral's run), Haiku 5.5 at 39.2% (Anthropic's run) and 33% (AA's run). Those three are comparable to each other. None of them is comparable to the Terminal-Bench v2.1 figures edition 018 carried.
Reported-but-unconfirmed, and a chain worth stating. AA's full comparison table reaches this brief through a third-party write-up of AA's article plus AA's own model page and posts; AA's article URL itself returned 404 at compile time at artificialanalysis.ai/articles/claude-haiku-5-5. The index figure of 43 and the #43-of-696 rank were read directly from AA's model page. The comparator scores are rounded in that write-up and do not match edition 019's unrounded figures to the decimal — Kimi K3 appears there as 44 against 43.6 previously, DeepSeek V4.1 Flash as 39 against 39. Treat the ordering as solid and the last digit as not.
Customer figures, which are the parties' own. Anthropic's post carries HubSpot at 92.8% on its own CRM suite averaged over three runs, AlphaSense at 0.84 against 0.76 for Haiku 4.5 across 400 queries, Box 11 points higher at about half the latency, Asana more than 30% lower latency, and Cognition's Devin Fusion at 66.2 on FrontierCode with Haiku 5.5 as sidekick. Every one is a customer's count on a private evaluation; none is independently run.
Sources Anthropic, Introducing Claude Haiku 5.5 (primary, 7 Oct) · Anthropic model documentation (1M context, 128K max output) · Artificial Analysis model page (index 43; rank #43/696; read at compile time) · officechai, 7 Oct (carrying AA's full table) · @ArtificialAnlys — read 8 Oct
OpenAI ships GPT‑6 to every ChatGPT user and, the same morning, publishes the evaluations showing safety regressions — including three for under‑18 accounts
Three OpenAI posts carry the same date, 7 October, and they have to be read together. GPT‑6 and Intelligent UI for everyone rolls GPT‑6 into ChatGPT globally — Plus, Pro, Business and Enterprise on 7 October running GPT‑6 Sol, Free and Go “the next day” running GPT‑6 Luna — against a stated base of over 1.2 billion weekly users. Helping teens learn, plan, and shape the future of AI adds a College Planner, flashcards and quizzes to ChatGPT for Teens. And the deployment-safety hub's GPT‑6 Sol and GPT‑6 Luna: October 2026 update reports what the evaluations found.
What they found, in OpenAI's own words and measured without system-level safeguards: GPT‑6 Sol shows a “statistically significant regression” against GPT‑5.6 on standard self-harm, and GPT‑6 Luna shows statistically significant regressions on standard self-harm, gore and sexual content. For under‑18 accounts, both October models regress against their August counterparts on age-restricted content, sexual content and emotional reliance, with Luna also regressing on gore. OpenAI states that manual review found the violations borderline and generally safe, that red teaming surfaced no high-severity risks, that a classifier-based block for self-harm, sexual content and gore operates as a system-level mitigation not captured in those scores, and that the emotional-reliance regression is partly an artefact of an evaluation oversensitive to benign nicknames. All of that is OpenAI's, published by OpenAI, on the day of the rollout.
Also on 7 October, Common Sense Media's Youth AI Safety Institute published an assessment rating ChatGPT for Teens “Unacceptable Risk” and saying OpenAI should restrict ChatGPT to users 18 and over until the gaps are closed. OpenAI disputes the methodology. Both positions are below.
| Domain | Classification | Note |
|---|---|---|
| Cybersecurity | High, below Critical | Sol comparable to GPT‑5.6 Sol, no clear capability improvement; Luna below Sol |
| Biological and chemical | High | Sol did not cross the indicative Critical thresholds; Luna not tested at Critical, having scored below Sol on every High evaluation |
| AI self-improvement | Below High | Neither model reaches the threshold |
Confirmed numbers from the safety update. Instruction-hierarchy adherence: Sol 99.99%, Luna 99.79%. ExploitBench at maximum reasoning effort: Sol 82.62%, Luna 44.66%, with OpenAI's own caveat that results may be inflated by exposure to historical vulnerabilities. SEC-Bench Pro: Sol 68.85%, Luna 48.50%, against 85.4% for GPT‑6 Astra — the consumer models are materially below the frontier model on the cyber harness. Unwanted persistence after barriers, at maximum reasoning effort: GPT‑5.6 Sol 34%, GPT‑5.6 Luna 27%, GPT‑6 Sol 28%, GPT‑6 Luna 15.9% — and OpenAI states the metric was redefined to track successful circumventions, so that row is not like-for-like.
Bio capability against the High thresholds, which is where the margins are thin. Multimodal Troubleshooting Virology (pass@1), threshold 31%: Sol 51.68%, Luna 48.13%. Tacit Knowledge and Troubleshooting (cons@32), threshold 80%: Sol 72.45%, Luna 94.00% — the smaller model over the threshold the larger one is under. TroubleshootingBench (pass@1), threshold 36.4%: Sol 41.83%, Luna 41.46%. ProtocolQA Open-Ended (pass@1), threshold 54%: Sol 25.62%, Luna 34.57%. At the Critical thresholds Sol is clear on all three reported: SHP2 protein-function prediction 0.23 against a 0.60 threshold, coronavirus–ACE2 cell-entry composite 0.349 against 0.75, phage–plasmid co-evolution 12.95 against a ≤9.40 threshold where lower is worse for safety.
What the independent assessment actually tested. Common Sense Media's Youth AI Safety Institute reports running more than 4,000 prompts on accounts registered to 13- to 17-year-olds, twice — before and after OpenAI announced ChatGPT for Teens on 18 August. Its findings: on more than a dozen newly created parent-linked accounts, testers spent up to an hour on suicidal ideation, self-harm or disordered eating with zero alerts reaching the linked parent account; the system missed more than one in four warranted crisis referrals and fell below its 95% threshold on three of five severe harms treated as Red Lines; study mode could be exited; and adult-registered test accounts never switched to the under‑18 experience across repeated testing. The Institute discloses that its funders include the OpenAI Foundation.
OpenAI's rebuttal, which is specific rather than general. Its spokesperson told Axios: “We welcome rigorous independent evaluation, but we do not believe Common Sense Media's testing accurately reflects” the product. On the alert finding, OpenAI says much of the testing ran before parent and teen account linking had finished, which can take several hours. On age estimation, it says prediction draws on multiple signals and can take up to two weeks. On crisis referrals, it says its larger-scale data shows more hotline resources displayed to under‑18 users over the period studied. On study mode, it confirms that exiting during Study Hours is intentional, designed as flexible rather than as a hard parental lock. Three of those four are testable claims about timing and sampling, and none has been independently checked at compile time.
The inference, and its premise. Read the regressions and the rollout as one decision rather than two: OpenAI measured the regressions, judged them acceptable on manual review and system-level mitigation, and shipped on the same day it published them. The load-bearing premise is that the safety update and the rollout post describe the same model build at the same moment, which both documents state. What this brief is not claiming: that the regressions caused any harm, that the classifier mitigation is inadequate, or anything about why the regressions appeared. No source read here addresses the cause.
Intelligent UI, briefly, because the capability claims are internal only. ChatGPT now composes answers from streamable interactive components — charts, forms, buttons, generated tools — rendered by a compiler as the model generates. The two speed figures are OpenAI's own: GPT‑6 Extra High “begins answering in the same amount of time as GPT‑5.6 Medium” while scoring better than GPT‑5.6 Extra High, and GPT‑6 Instant “starts answering 44% sooner, on average, than GPT‑5.6 Instant” on web-search questions. No independent run of either was found at compile time at artificialanalysis.ai. No API release and no pricing accompany this rollout.
Sources OpenAI, GPT‑6 and Intelligent UI for everyone (primary, 7 Oct) · OpenAI Deployment Safety Hub, GPT‑6 Sol and GPT‑6 Luna: October 2026 update (primary, 7 Oct) · OpenAI, Helping teens learn, plan, and shape the future of AI (primary, 7 Oct) · Common Sense Media Youth AI Safety Institute (primary, 7 Oct) · Axios, 7 Oct (OpenAI's response)
Epoch puts a ratio on the chip-exposure question — China 2.7× the US — and a second number on how far behind Chinese AI revenue actually is
Two Epoch AI reports, both dated 6 October and neither carried in edition 019. The first, by Daniel Carey, separates semiconductors out of the broader electronics category in OECD input-output tables and traces chips to the spending that pays for them. The headline: China is roughly 2.7× as exposed to semiconductor supply shocks as the United States. In 2022 every $1,000 of Chinese final demand generated $15.2 of revenue for semiconductor producers, against $5.7 for the same US spending — because, as the authors put it, “semiconductors account for a larger share of the cost of the goods it buys.”
The second, by Cheryl Wu and Anson Ho, estimates that “China's six leading AI firms collectively earn about 10% of OpenAI and Anthropic combined, in terms of AI-related revenue.”
| Company | Run rate | Note |
|---|---|---|
| Anthropic | 65 | US |
| OpenAI | 40 | US |
| ByteDance | 4.0 | Model-as-a-Service revenue only; the ~$4bn target may include enterprise sales |
| Alibaba | 2.4 | Model-as-a-Service revenue only |
| Z.ai | 1.8 | Appendix figure annualised from an interim report using 2025 seasonality |
| Moonshot | 1.0 | API sales dominate |
| DeepSeek | 1.0 | API sales dominate |
| MiniMax | 0.8 | Appendix figure annualised from an interim report |
What the exposure model actually simulates. Epoch calibrates a trade model on the revised tables and runs two shocks together: a large drop in Taiwanese productivity, and a China–West decoupling. Two to three years out, the combined shock raises Chinese advanced-processor prices roughly 17-fold against roughly 20% for the US, and cuts Chinese real gross national expenditure about 3% against about 0.6%. Exposure by producing economy, China relative to the US: 4.7× for Taiwan, 4.0× for South Korea, 3.6× for Japan. Epoch reports the ordering as robust to parameter changes in every case tested.
The caveats Epoch prints itself, which matter for how far this travels. The underlying data end in 2022 — before the most recent rounds of export controls and before most of the domestic-substitution push that the question is usually asked about. Epoch also states that its aggregation likely overstates China's exposure, and that splitting Chinese manufacturing narrows the gap without closing it. A modelled ratio on four-year-old input-output tables is a useful order of magnitude and not a forecast, and this brief is not treating it as one.
Why the revenue table is harder to use than it looks. The run rates come from different months spanning June to September 2026; the Alibaba and ByteDance figures capture Model-as-a-Service revenue only, which Epoch says understates firms that likely monetise indirectly through cloud and advertising; two of the Chinese figures are annualised from interim reports using 2025 seasonality and differ from management-reported numbers in Epoch's own chart. These are revenue estimates and nothing else — no source read here says anything about costs, margins or profitability at any of these companies, and this brief is making no such claim.
Sources Epoch AI, Who is most exposed to a chip supply shock? (6 Oct) · Epoch AI, How do Chinese AI companies make money? (6 Oct) · Epoch AI publications index (dating, read at compile time)
Finland orders work stopped at two Google data-centre sites over missing environmental assessments
Reported 7 October. The Finnish supervisory agency responsible for environmental matters ordered Tuike Finland, the company representing Google, to suspend work at data-centre sites in Muhos and Kajaani until mandatory environmental impact assessments are completed. The agency is investigating whether hundreds of hectares of forest were cleared at the two sites without the assessment; a Finnish conservation group puts the figure at more than 300 hectares. Two dates are attached: the company must explain its plans by 14 October or face enforcement proceedings, and must halt preparatory measures that would significantly alter the environment no later than 23 October. Google announced a €13 billion ($14.6 billion) Finnish digital-infrastructure programme in September; a Google spokesman said the company would “study the LVV's findings and follow their guidance.”
What this is and is not evidence of. It is a procedural order about environmental-assessment sequencing at two named sites, with named deadlines. It is not a cancellation, not a finding of liability, and not a statement about capacity: no source read here gives a megawatt figure, a build schedule, or what the sites were to be used for. Set against the 3.6 GW of PJM-area power Google contracted on 6 October (edition 019), the pattern worth watching is that the binding constraints on data-centre buildout are showing up as permitting and grid-interconnection questions in specific jurisdictions rather than as chip supply — that is this brief's reading across two items, not a claim either source makes.
Sources Al Jazeera, 7 Oct (the order, the deadlines, Google's statement) · TechTarget, 7 Oct (independent dating; the €13bn figure) · Quartz, 7 Oct
Also on the wire
Confirmed, and not enough on their own to move the picture.
-
Google ships EmbeddingGemma 2 under Apache 2.0 — 740M parameters, five modalities, weights out
Posted 6 October by two Google DeepMind research engineers. Text, code, image, video and audio into one embedding space; text-only use needs as little as 270M parameters, with optional vision (170M) and audio (300M) encoders. Context 8K tokens, four times EmbeddingGemma 1, covering about 5.5 minutes of audio or 29 images. Output 768 dimensions, truncatable to 512, 256 or 128 via Matryoshka representation learning. On a Pixel 11 Pro with quantisation: about 191MB active RAM text-only, about 567MB full multimodal. The one hard benchmark figure in the post is MTEB Code, 68.76 to 78.68; the image and audio charts carry no values in the text. Weights on Hugging Face and Kaggle. Unlike most items in this brief's recent weeks, the weights actually shipped with the announcement.
Source blog.google, 6 Oct (primary) · @sundarpichai — read 8 Oct from the Frontier Wire Sources list
-
Artificial Analysis starts labelling derived models on its video boards
Announced on the list within the window: AA's video leaderboards now label models post-trained from another model on the same board, and let readers hide them. AA's stated example is that two of the top five on AA-Video-T2V v2.0, Utopai X and MiniMax H3 Max, are built on MiniMax H3. Separately AA placed Vidu Q4 Preview at #3 on AA-Video-I2V v1.0, from #19 for Vidu Q3 Pro. The labelling change is the more durable item: a leaderboard whose top five contains three variants of one base model measures something different from what a reader assumes it measures.
Source @ArtificialAnlys — read 8 Oct from the Frontier Wire Sources list · AA-Video-I2V v1.0 leaderboard
-
A subsidised compute pool for .edu and .gov goes to public-sector access
Circulated on the list: National Compute, which describes itself as “pooled infrastructure for the frontier research ecosystem,” is taking sign-ins from .edu, .mil and .gov addresses, with a waitlist for everyone else; the site states “the Grid is only available to public access partners during preview.” The post carrying it claims hundreds of MI355x and B300 nodes are live and subsidised. That hardware claim is the poster's, not the site's: no node count, no hardware specification, no pricing and no launch date were found at compile time at nationalcompute.com, which names only a capacity tracker and an evaluation environment. Carried because compute access rather than funding is the stated bottleneck for academic work — a point made on the list by Nathan Lambert — and because a subsidised public-sector pool at that scale would matter if the specification holds up.
Source nationalcompute.com (read at compile time) · @natolambert, @AnjneyMidha — read 8 Oct from the Frontier Wire Sources list
-
Day four, and the Super Intelligence Force charter is still not found at compile time
Checked again this morning. At compile time on 8 October, the first page of whitehouse.gov/presidential-actions, which is ordered newest first, lists four October items: National Energy Dominance Month, 2026 and Columbus Day, 2026, both dated 7 October, the diesel-fuel emergency tax relief order of 5 October, and the National Manufacturing Day proclamation of 2 October. Nothing dated 6 or 8 October appears, and no action establishing or chartering a Super Intelligence Force was found there. The renaming order of 29 September, Inaugurating the Era of Super Intelligence, is still a different document. Ninety-six hours after the announcement carried in edition 017, the charter remains unreadable. The two dated deliverables are unchanged: a 120-day report falling at the start of February 2027 if the clock runs from 4 October, and the renaming order's 60-day legislative-language obligation falling at the end of November.
Source whitehouse.gov/presidential-actions (first page, read at compile time)
-
The lab desks, checked
Checked at compile time. Anthropic's newsroom carries Haiku 5.5 (item 02) and nothing dated 8 October; its research index carries nothing newer than 1 October. OpenAI's newsroom carries the three 7 October posts in item 03 and nothing dated 8 October. Google: EmbeddingGemma 2 ran at blog.google on 6 October, which is where this brief found it; deepmind.google's blog index gives month-level dates only and nothing could be dated to 7 or 8 October from it, for the thirteenth consecutive edition. METR carries nothing newer than the 6 October observability post in edition 019. ARC Prize's blog is unchanged since 3 September. Epoch AI's most recent publications are the two 6 October reports in item 04; its capabilities data hub shows a 7 October refresh, and a refresh is not a publication. Mistral carries nothing after the 6 October Large 4 post, and the weights promised for “end of this month” have not appeared on that page. x.ai's news index returned nothing newer than 28 September at the URL tried; Meta's AI blog nothing since July. Qwen's old blog index now says the blog has moved and carries nothing newer than September 2025; nothing could be dated at qwen.ai/research at compile time. DeepSeek's news page did not render a dated item list to this brief, and its most recent linked entry remains the 10 September one.
Source Anthropic newsroom · Anthropic research · OpenAI · blog.google AI · DeepMind blog index · METR · ARC Prize · Epoch AI · Mistral · x.ai · Meta AI · Qwen · DeepSeek (all read at compile time)
Checked and spiked
Circulating this morning; did not survive checking.
“Terence Tao's statement” on the OpenAI release. At least one aggregator is carrying the 7 October statement as Tao's own. It is not. The document is issued by the Communications Working Group of the Association for Human Mathematics, lists no individual signatories, and appears on Tao's blog as a guest post reproduced from AHM's own statements page. Hosting a statement is not writing one, and the difference matters precisely because the statement is a call to collective action whose weight depends on who is behind it — which, on the document as published, is an organisation's working group and not a named person. This is the second time a guest post on that blog has been attributed to its host: edition 007 recorded the same thing happening to Po-Shen Loh's 19 September piece. Carried in item 01 with the attribution the document itself carries.
A joint declaration by 25 Fields Medalists, recirculating as new. Posted to Mathstodon on 11 September 2026 — four weeks before the OpenAI release it is now being read as a response to. It is a separate document with separate signatories and this brief is not treating it as part of this week's reaction. Edition 019 spiked the same document on the same ground twenty-four hours ago; it is recirculating faster now that there is a real statement for it to be confused with. Anyone citing it against the 6 October release should check the date first.
A discounted price for Mistral Large 4. A model-tracking site lists Mistral Large 4 at a “sale” price of $0.68 / $2.09 per million tokens against a $1.36 / $4.18 list, with no end date. Edition 019 carried the $1.36 / $4.18 figures from Mistral's own post. No discount, promotional price or end date was found at compile time at mistral.ai/news, and the brief is not carrying a price it cannot trace to the seller. The same listing gives Large 4 as 1.05T total and 52B active parameters against the 1T / 49B in Mistral's post, and a DeepSWE figure that Mistral's post does not contain. Unresolved, so not carried.
Corrections
Errors in this brief — fixed in place above, logged here.
No new corrections. Nothing in editions 001–019 has been flagged by a reader or found in error since edition 019 went out. Editions 002, 003, 006, 008, 014 and 015 carry their own corrections, archived below with their editions — edition 015's correction, logged in edition 016, is archived with that edition. The standing note, carried forward: a statement that was true when written and overtaken by a publication hours later is not an error, and this brief will keep saying so rather than quietly retrofitting; the test is whether it is applied honestly in both directions. The rule adopted after edition 015 — that an absence is reported as “not found at compile time, at this URL” or it is not reported — is applied seven times in this edition: to peer review and to any advisory-group response in item 01, to an independent run of OpenAI's two speed figures in item 03, to AA's own article URL in item 02, to the hardware specification behind the compute pool and to the presidential action in Also on the wire, and to the lab desks above.