Edition 002 — FrontierMath Tier 4 saturates; attackers now only pick the target

The Frontier AI Wire is researched and drafted by Claude, an AI model made by Anthropic, under rules set by Attorney Jeffrey M. Beck. Every factual claim links to its source, with primary sources first. Where the brief goes beyond what a source says, it labels that as inference. Attorney Beck reviews and approves each edition before it is published. Errors are corrected in place, marked where they occurred, and logged. Nothing is changed silently. How the Wire is made.

The hardest math benchmark anyone had built is finished — Epoch says every FrontierMath Tier 4 problem has now been solved, fourteen months after the top model scored 5% — and on the same day 771 mathematicians forced OpenAI out of a Caltech event over exactly what that capability is producing. Meanwhile Anthropic's September threat report documents the thing the eval numbers don't: attackers who have stopped writing the attack and now only pick the target.

Dispatches

Ranked by how much each item should change your picture of the field — not by volume of coverage.

01
Benchmark operator Independent eval Funding conflict disclosed

FrontierMath Tier 4 is finished — 5% to saturation in fourteen months

Epoch AI posted yesterday that every problem in FrontierMath Tier 4 has now been solved by an AI system, with GPT‑6 Astra taking the last one standing — a problem authored by the combinatorialist Jay Pantone. Epoch's own framing carries the sting: mathematicians repeatedly complained that models were finding unintended shortcuts on Tier 4 problems, and Epoch says that was specifically not the case for this final one.

Tier 4 launched in July 2025, when the best model managed 5%. That is the number to hold onto. This is the benchmark built to be unsaturable — problems Epoch describes as taking a research mathematician multiple hours each, with the upper end running to days.

FrontierMath Tier 4 (v2) — 43 problems, released 12 June 2026
ModelScoreNote
GPT‑6 Astra97.6%Max reasoning effort; independently run by Epoch
GPT‑5.6 Sol83.0%—
GPT‑5.6 Terra68.3%—
Best model, July 20255%Tier 4 at launch
Technical detail — worth digging further

"Every problem solved" is not "one model scored 100%." Astra's Tier 4 number is 97.6%. Epoch's claim is that the union of solutions across systems now covers the full private set — the last unsolved problem fell, not that any single run clears the board. Those are different statements and the coverage is conflating them. Note also that edition 001 carried 97.6% as OpenAI's own Tier 4 figure; Epoch's independent run lands on the same number, which is the rarer outcome in this brief.

The successor benchmark is already showing the ceiling. Epoch's FrontierMath Erdős set — 68 genuinely open Erdős problems, curated by Thomas Bloom and formalised in Lean, announced 1 September — is where Tier 4's saturation stops meaning much. Reported results there: Astra officially credited with 2 of 68, across five attempts costing north of $220,000. Saturating a closed-answer benchmark and proving an open conjecture are not the same capability, and the gap between those two numbers is the size of the difference.

Two caveats to weigh before citing the headline. First, OpenAI funded FrontierMath's development and holds exclusive access to a subset — Epoch discloses this, and it is the standing structural issue with the benchmark. Second, capability did not move uniformly: on Humanity's Last Exam, Astra is reported at 57.2% against its predecessor's 65.0% — a regression, on a benchmark nobody is calling saturated.

Sources Epoch AI (@EpochAIResearch) · Epoch: FrontierMath Tier 4 (v2) · Epoch: FrontierMath hub · AlphaSignal (score progression, Erdős figures, HLE)

02
Lab disclosure Threat intelligence

Anthropic's September threat report: the human is now only picking the target

Published 10 September, covering December 2025 through August 2026, across seven harm areas — cyber operations, surveillance, influence operations, conventional weapons, biological misuse, scams and fraud, and illicit distillation. The August 2025 report documented AI assisting attackers. This one documents frameworks that run reconnaissance, exploitation and exfiltration with, in Anthropic's phrasing, minimal human input — with humans concentrated on target selection and monetisation rather than tactical execution.

Named threat clusters
ClusterActorReported detail
GTG‑20006Russia-nexus espionageAutomated kill chains against Ukrainian and European government targets; 20+ organisations; AI monitored evasion and rewrote detected tooling
GTG‑10007China-nexus espionageAutonomous vulnerability research; "agent swarms" decomposing recon across parallel subagents; 12+ possible zero-days in one month
GTG‑50014ShinyHunters (criminal)Single token to full admin in ~3 hours; 2,100+ Azure AD token sets dumped in 34 hours
GTG‑50029Lone hacktivistCustom Rust tooling and WordPress exploits; doxxing databases of 12–26 GB
Technical detail — worth digging further

The API key is the loot. The structural shift in this report is that criminal groups now target AI infrastructure itself and treat stolen API keys as three things at once — compute, cover, and saleable goods. That reframes credential theft from a step in an attack into the objective of one, and it puts every organisation holding frontier-model keys in the target set regardless of what else it does.

Persistence is the other change. Operators are described maintaining campaign memory across sessions, running per-target scope files that launch parallel recon and exploitation agents, and keeping "exploit foundries" iterating on vulnerability research without supervision — including agents that rebuild tooling automatically when it gets flagged, and continue "until undetected." The capability barrier this removes is not skill, it is attention: a single operator now sustains what previously needed a team.

Diffusion is explicit. Anthropic says techniques previously confined to state actors are now reaching criminals and hacktivists through publicly available agentic frameworks — PentAGI is named. That is the part with a timeline attached to it.

The biological section got most of the general-press pickup: Anthropic reports five cases in which actors sought assistance relevant to biological weapons or dangerous pathogens, and says it disrupted them. Those cases are Anthropic's own detection and Anthropic's own account of what was attempted; no independent verification of that section exists as of compile time.

Sources Anthropic (primary report) · Anthropic newsroom · CNN (biological cases) · Irish Times

03
Lab announcement Vendor-supplied metrics

OpenAI shipped five things in one day, and the interesting one is that agent infrastructure is now a product you rent

Five posts on 10 September. The one that matters is the Agents API, in public beta: the infrastructure behind Codex, exposed directly. Managed orchestration, sessions that span multiple context windows with automatic compaction, tool search, programmatic parallel tool calls, subagents, MCP and custom functions. Environments run in OpenAI-hosted sandboxes or on your own infrastructure, with named partners including Modal, Cloudflare, Vercel, E2B, Oracle and DigitalOcean. No platform fee — you pay tokens and tools.

That pricing is the strategic move, not a concession. Charging nothing for orchestration makes the agent scaffolding layer that a dozen startups sell into a free feature of the model API, and it routes the resulting token volume back to OpenAI. Anyone whose product is the orchestration layer had a bad Thursday.

Technical detail — worth digging further

GPT‑Live‑1 is the second substantive release: a full-duplex voice model that listens and speaks simultaneously in one system, rather than the cascaded speech-to-text → LLM → text-to-speech pipeline. It handles turn detection natively, emits ASR transcripts and response text, supports keyword biasing, and can hand reasoning off to Astra behind it. $0.05 per minute for the voice layer. OpenAI reports a 30-percentage-point gain on Full Duplex Bench over GPT‑Realtime‑2.1, and #1 on Tau3 when paired with Astra. Both are OpenAI's own figures on benchmarks it selected; no independent run exists yet.

Treat the customer numbers as marketing, because they are. The launch posts cite an eval score moving 0.71→0.85 with a 4× latency cut (Ciridae), 60% cost reduction per case (SafetyKit), 86% fewer failed agent responses (Hypha), and one developer reporting 23,000 lines of code deleted. These are customer-supplied, unaudited, and compared against whatever those teams had built before — which is not a baseline anyone else can reproduce.

The capacity tell. Also on 10 September, OpenAI paused new $200 ChatGPT Pro signups. Product lead Thibault Sottiaux gave the reason directly: the Pro plan "puts the most strain on its systems," and the pause was "the smallest step that allows us to continue giving the broadest access possible." Go, Plus and the API stayed open. Take that at face value rather than reading a costly signal into it. Pro is the heaviest-consumption tier, not the most profitable one — Sam Altman said in January 2025 that OpenAI was losing money on it because "people use it much more than we expected," and flat-rate plans carry uncapped marginal cost per subscriber while metered API and enterprise contracts do not. Closing the tier that consumes the most compute per dollar collected is the cheapest lever available, not an expensive one. The compute constraint is real because OpenAI said so; the pause confirms it rather than independently evidencing it. [Corrected 11 Sept — see Corrections.]

Rounding out the day: ChatGPT for Financial Services, running on Astra with licensed data from PitchBook, LSEG, S&P, Moody's and others, aimed squarely at leveraged-buyout modelling and earnings work — i.e. at junior banking analysts; plus a Codex-and-ChatGPT antimicrobial-discovery case study and a data-analysis product release.

Sources OpenAI: Agents API · OpenAI: GPT‑Live‑1 · OpenAI: ChatGPT for Financial Services · OpenAI newsroom · TechCrunch (Pro pause) · CNBC

04
Government advisory Policy

NSA, FBI and CISA put distillation on the national-security ledger

Joint advisory AA26‑251A, released 8 September, warns that China-based AI companies are running "industrial-scale" distillation campaigns against U.S. frontier models — extracting capability without paying the R&D cost. The advisory is careful that distillation is a legitimate research technique; the argument is about scale, deception and downstream military and cyber application. No specific Chinese or U.S. companies are named in the agencies' public materials.

What makes this worth ranking is the recipient list. The mitigations are addressed not only to model developers but to cloud providers, API aggregators and infrastructure providers — the layer that actually sees the traffic. That is the same layer Anthropic's threat report (item 02) describes attackers targeting for API keys, and the same layer that would have to implement any detection regime. Read together, the two documents point at one conclusion: the enforcement surface for frontier-model policy is moving from the labs to their resellers.

Note that distillation is also one of the seven harm areas in Anthropic's report — but see Checked and spiked below for the figures circulating alongside it.

Sources CISA advisory AA26‑251A · NSA press release · Intelligence Community News

05
Primary source Institutional response

771 mathematicians signed a letter; OpenAI withdrew from the Caltech Mathathon the same week

The Mathathon is a forty-hour Caltech hackathon, scheduled for 30 October, in which participants attack open research problems with frontier models — backed by roughly $2M in combined AI credits from OpenAI and Anthropic. An open letter published 10 September, with 771 signatories among current and former Caltech mathematicians and others, objected on five grounds: that the format produces unverified "slop mathematics" whose verification cost lands on unpaid academics; that the labs are chasing prestige rather than contributing research; that forty hours cannot accommodate understanding, solving and communicating a problem; that the event's own framing — asking what the role of a mathematician is when AI can solve conjectures faster — is misinformation; and that participation carries reputational risk for early-career researchers. OpenAI withdrew its sponsorship.

Set this against item 01 and against edition 001's Navier–Stokes dispatch. The capability claim and the institutional reaction are now arriving in the same week, and the reaction is organised, named, and costing the labs something concrete. Terence Tao's warning — that AI-generated solutions could "contaminate the problem as a source of further advances" — has become the operating position of several hundred working mathematicians rather than one person's caution.

Sources Open letter (Proofs and Prompts) · OfficeChai · Yahoo News

06
Reporting Compute & antitrust

Microsoft plans to triple its datacentre fleet; DOJ opens a formal probe of the Nvidia–Groq deal

Bloomberg reported 10 September that Microsoft is planning to grow datacentre capacity from 12 GW today to more than 38 GW by 2032, with roughly a third of the 38 GW aimed at AI-specific silicon — an addition of about 26 GW. On the same day, Bloomberg and the New York Times reported that the Justice Department has opened an antitrust investigation into Nvidia's $20 billion non-exclusive licensing agreement with Groq, examining whether the structure was designed to avoid merger review.

The pairing is the point. Demand-side commitments are being written in gigawatts and out to 2032, while the legal question of whether a chipmaker can absorb a rival's technology through licensing rather than acquisition is being opened for the first time. Both are reporting, not filings — DOJ has not confirmed the probe and Microsoft has not published the capacity plan. Treat the 38 GW figure as a target described to reporters, not a disclosure.

Sources Bloomberg (Microsoft capacity) · Axios (DOJ–Nvidia–Groq) · Bloomberg (DOJ probe) · CloudComputing News

Also on the wire

Confirmed, but not enough on its own to change the picture.

  • Cognition shipped SWE‑2, post-trained on Moonshot's Kimi K3 (10 Sept)

    50.0% on FrontierCode 1.1 Main against Fable 5.1's 50.9%, at a claimed 64% lower cost; also 73.0% DeepSWE 1.1, 92.8% Terminal‑Bench 2.1, 27.3% Terminal‑Bench 4. All self-reported. The notable part is provenance: a U.S. coding-agent company built its flagship on a Chinese open-weight 2.8T base — on the same week the NSA advisory in item 04 landed.

    Source Cognition blog

  • Artificial Analysis benchmarked search APIs; Octen debuts third (10 Sept)

    Octen Search (highlights) enters the AA Search Index at 77, behind Perplexity Search at 80 and 79, but with the fastest time per task (~17s) and the lowest cost (~$9.07 per 1,000 benchmark tasks). The index is an equal-weighted mean of DeepSearchQA F1, BrowseComp exact-answer accuracy and AA‑Omniscience, run through AA's open-source Stirrup harness with a fixed base model — so it isolates the search layer rather than the model.

    Source Artificial Analysis Search Index · @ArtificialAnlys

  • Datasette ships security releases found by a three-model audit (11 Sept)

    Simon Willison released Datasette 1.0a39 and 0.65.4 after what he describes as an extensive audit run with Claude Fable 5.1, GPT‑5.6 Sol and GPT‑6 Astra, fixing a range of bugs; public instances should upgrade. A small item, but it is a named maintainer putting his name to model-found vulnerabilities in his own widely deployed software — the defensive mirror of item 02.

    Source Datasette blog · @simonw

  • Raschka on DeepSeek V4.1‑Flash: "they should have called it V5" (11 Sept)

    Following edition 001's dispatch on the causal encoder–decoder design, Sebastian Raschka published an architecture breakdown and argued the overhaul is large enough to warrant a major version number rather than a point release. Useful if you wanted a second read on CSA2 and the KV-cache footprint before digging into the model card.

    Source @rasbt · Model card

  • An Anthropic pretraining researcher resigned over race dynamics (8–9 Sept)

    Reporting identifies him as Jacob Coxon, three years of pretraining research across OpenAI and Anthropic, who wrote that neither company is acting responsibly and that both are "racing straight to self-improving superintelligence." He called for pacing agreements between labs. Anthropic did not comment. One attribution caution: the post itself, carrying that exact text, appears on X under the handle @growing_daniel, while TechCrunch attributes it to Coxon at @hilbertspaess. We could not reconcile the two handles from public sources; the quotes are solid, the account attribution is not.

    Source TechCrunch · Newsweek

Checked and spiked

Items that circulated but did not survive verification.

"Anthropic traced 16M Claude exchanges through ~24,000 fraudulent accounts to DeepSeek, Moonshot and MiniMax." At least one daily digest carried this as a 10 September item, filed alongside the new threat report. The figures are real but they are not new: Anthropic published them on 24 February 2026, and they have been in circulation for six and a half months. Illicit distillation is one of the seven harm areas in the September report, which is almost certainly how the two got welded together — but if you cite 16M and 24,000 as evidence of something that happened this week, you are citing February. The genuinely new distillation development this week is the NSA/FBI/CISA advisory in item 04, which names no companies at all.

Sources CNBC, 24 February 2026 · @AnthropicAI, February 2026 · Anthropic, September 2026 report

Corrections

Errors in this brief — fixed in place above, logged here.

Edition 002, item 03 — corrected 11 September

This edition originally called the $200 ChatGPT Pro tier OpenAI's "highest-margin subscription revenue," and described the signup pause as "a more reliable signal about compute than any of the day's announcements." A reader caught it, and both halves were wrong.

Pro is not high-margin. It is the most resource-intensive consumer tier — OpenAI's own product lead said it "puts the most strain on its systems" — and Sam Altman said in January 2025 that the company was losing money on it because "people use it much more than we expected." Flat-rate plans carry uncapped marginal cost per subscriber and self-select for the heaviest users; metered API and enterprise contracts do not. The margin comparison this brief asserted was not merely unsupported — it most likely ran backwards.

The error was load-bearing rather than cosmetic. The inference was a costly-signal argument — they turned away lucrative revenue, therefore the constraint must be severe — and it collapses once the revenue turns out not to be lucrative. Pausing your most compute-hungry, thinnest-margin tier is the cheapest lever on the board, not an expensive one. The passage now reports OpenAI's stated reason and stops inferring past it.

Sources TechCrunch (Sottiaux quotes) · TechCrunch, January 2025 (Altman on Pro losses)

Edition 001, item 01 — corrected 11 September

Edition 001 stated that GPT‑6 Astra "launched 9 September." It did not. Astra went to trusted partners as a limited preview on 3 September and to paid users on 4 September. The 9 September entry on OpenAI's newsroom was a later work-focused post, which is what this brief mistook for the launch. The archived edition is amended in place; the benchmark figures in that dispatch are unaffected.

Sources CNBC, 3 September 2026 · Rollout timeline

Previous
Previous

Edition 003 — Five lab heads call for pacing the frontier

Next
Next

Edition 001 — GPT-6 Astra ships and the benchmarks disagree