Edition 005 — Meta answers on pacing; Google's voice model takes the index
The Frontier AI Wire is researched and drafted by Claude, an AI model made by Anthropic, under rules set by Attorney Jeffrey M. Beck. Every factual claim links to its source, with primary sources first. Where the brief goes beyond what a source says, it labels that as inference. Attorney Beck reviews and approves each edition before it is published. Errors are corrected in place, marked where they occurred, and logged. Nothing is changed silently. How the Wire is made.
Meta was the blank row in the pacing table, and filling it in showed the coalition has a second half. Zuckerberg posted a point-by-point case that labs need no pact to build safely, Meta's superintelligence chief called recursive self-improvement “one of the riskiest pathways,” Altman told a San Francisco audience he is “confident” the industry can police itself, and Nvidia's Huang called the antitrust waiver behind Amodei's step two “completely unnecessary,” per CNBC. Underneath the argument, the actual shipping: Google put a voice model into production that beats OpenAI's five-day-old one on the independent index and costs 40% less to run.
Dispatches
Ranked by how much each item should change your picture of the field — not by volume of coverage.
Meta answered on pacing, and the answer separates evaluators from coordination
Edition 003's table carried one blank row: Meta, no public statement. It is filled in. On 15 September Mark Zuckerberg posted a numbered case on X — 3.9M views, reposted by AI at Meta — that frontier labs already have sufficient incentive to build safely without an agreement among them. The load-bearing sentence: “Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens.”
What makes it more than a refusal is that he conceded two of Amodei's three steps while rejecting the third. On evaluators: “Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas … Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators.” On the rate of progress, he offered something none of the other principals has: “Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well.” What he rejected is the coordination: “Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would.”
Meta Superintelligence Labs chief Alexandr Wang pushed in the same direction on the same day, calling for “a strong governance framework across training and deployment” with external oversight of safety criteria, and naming the pursuit of recursive self-improvement “one of the riskiest pathways for potential loss of control to powerful models.”
| Principal | Embedded / independent evaluators | Coordinated rate limits | Antitrust cover |
|---|---|---|---|
| Amodei (Anthropic) | Yes — terms published, doing it unilaterally | Yes — the core ask | Says government mediation is needed |
| Altman (OpenAI) | “We will do the same”; no terms published | In talks on a standards body | Lehane, 15 Sept: no waiver needed |
| Zuckerberg / Wang (Meta) | Yes — says MSL already does it | No — unilateral only | Not raised |
| Hassabis (Google DeepMind) | Endorsed direction | Points to its standards-body proposal | Not stated |
| Nadella (Microsoft) | Endorsed; Code of Conduct published 14 Sept | Not specified in the document | Not stated |
| Huang (Nvidia) | — | Opposed | Waiver “completely unnecessary” (CNBC, 15 Sept) |
| Sacks (White House) | — | Unilateral only | No antitrust help for joint action |
The compute commitment is the one testable thing in the post. “The significant majority of compute towards serving people rather than racing towards recursive self-improvement” is a claim about an allocation ratio between inference-for-users and self-improvement research. Meta has published no figure, no definition of which workloads fall on which side of that line, and no reporting mechanism. It is nonetheless the first pacing-adjacent commitment anyone has framed in units that could in principle be audited — compute share — rather than in intentions. [Corrected 21 Sept — Anthropic published a three-part measurement schema with its own readings on 17 September; see edition 007 Corrections.] Amodei's step two asked for “limits on the rate of unchecked AI progress” and has never said what the rate would be measured in. Zuckerberg has now proposed a unit while declining the coordination. Those are separable, and worth tracking separately.
What Altman actually said, and where it sits. At a San Francisco conference on 15 September (reported by Bloomberg that day, by the Irish Times on the 16th): “I am very confident in our company's ability – our industry's ability – to do this safely,” adding that he was “disappointed by how it's been framed.” Read against Altman's own 12 September position — that OpenAI “will do the same” on evaluator access — this is not a reversal, but it is a different emphasis: four days earlier the argument was for outside eyes, and on Monday it was for trusting the industry. Both can be held at once. The brief notes the shift without claiming to know what drove it; no source establishes that.
The reading this brief offers, marked as a reading: last week's story was told as convergence, and edition 004 corrected that framing once already. The more durable structure now visible is a split along a different seam than “safety versus speed.” Independent evaluators have near-unanimous verbal support — Anthropic, OpenAI, Microsoft and now Meta — and nobody but Anthropic has published terms. Coordinated rate limits have exactly one advocate among the principals and three explicit opponents, counting the White House. The load-bearing premise here is that evaluator endorsement and pacing endorsement are different commitments; that premise is supported by the texts themselves, which take them up separately. What follows from it is that the standards body, if it arrives, is far likelier to publish standards than to set a rate.
Sources Zuckerberg on X (primary) · Forbes (Zuckerberg and Huang) · Asianet (Alexandr Wang quotes) · CNBC (Huang: “completely unnecessary”) · Bloomberg (Altman) · Irish Times (Altman, verbatim) · Axios on the August essay he cites
Google shipped a voice model that takes the independent index off GPT‑Live‑1 five days after it launched — at 40% of the cost
Released 15 September: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, native audio-in/audio-out models for real-time conversation. The unusual part is not the models. It is that Google's launch post leads with Artificial Analysis's numbers rather than its own, and AA's public leaderboard carries them — so for once the headline figure and the independently run figure are the same figure.
| Model (config as listed) | Index | Cost / hour |
|---|---|---|
| Gemini 3.8 Live Extended Thinking (High) | 82.6 | $3.50 |
| GPT‑Live‑1 (Astra, medium) | 81.5 | $5.83 |
| Grok Voice Think Fast 2.0 High | 81.3 | $4.80 |
| Gemini 3.8 Live | 76.0 | $0.84 |
| GPT‑Realtime‑2.1 High | 73.9 | $10.75 |
| GPT‑Realtime‑2 (High) | 73.6 | $4.14 |
Read the config column before citing the #1. AA's index is a weighted average of speech reasoning, agentic performance, arena preference and task success rate. The top Gemini entry is run at High; the top OpenAI entry is GPT‑Live‑1 (Astra, medium). A 1.1-point gap between a High configuration and a medium one is not the same claim as a 1.1-point gap between two models. No GPT‑Live‑1 High row appeared on the leaderboard at compile time, and this brief is not going to assume one exists or what it would score. The cost column is the more robust comparison, because AA measures it the same way for every row: $3.50/hr against $5.83/hr, and $0.84/hr for plain 3.8 Live at 76.0. Note this is AA's own cost-per-hour on the Big Bench Audio subset, not a list price — edition 002 carried OpenAI's $0.05/minute figure for the GPT‑Live‑1 voice layer, which is a different quantity and should not be differenced against these.
Google's own reported figures, separately. τ‑Voice agentic task completion 68.6%; Sierra's τ‑Voice‑banking 35.1%; Big Bench Audio reasoning 97.7%; 3.8 Live second on Speech Agent Arena. Sierra's leaderboard is third-party; the rest come through Google's post.
From the model card, which is worth reading against the launch post. 128K-token context in, 64K out; audio, image, video and text in; audio and text out; automatic detection and mid-sentence switching across 97 languages; background tool execution that does not interrupt the conversation. Two items to weigh: the knowledge cutoff is January 2025, and the frontier safety assessment was based on Gemini 3.7 Flash testing rather than a fresh evaluation — Google's determination is that the models reach no Tracked or Critical Capability Level and pose no material new risk over their predecessors. That is a reasoned inheritance, not an independent result, and the card says so.
Availability from day one: Gemini API and AI Studio for developers, Vertex AI, private preview in Gemini Enterprise, Search Live for 3.8 Live, and Gemini Live plus Gmail, Docs and Keep for Extended Thinking. The competitive fact worth holding is the interval: OpenAI launched GPT‑Live‑1 on 10 September and was off the top of the independent index by the 15th.
Sources Google (launch post) · Gemini 3.8 Audio model card · Artificial Analysis Speech to Speech Index · @OfficialLoganK (launch thread) · OfficeChai (cost comparison) · OpenAI: GPT‑Live‑1 (10 Sept)
Somebody killed the public price signal for AI compute. Commerce says it wasn't them.
Semafor reported on 15 September that the Commerce Department ordered the prediction market Kalshi to take down its AI-compute futures curve — an aggregate of several markets tracking the rental price of Nvidia chips and where it is headed — citing national security concerns, and that Kalshi complied. Per the reporting, Commerce also pushed the CFTC to freeze approval of new compute contracts for 60 days. The individual underlying markets remain live; the aggregated curve is what came down.
Two things have to be said plainly. First, the date: the order was given last month, not this week — 15 September is the date of the disclosure, and at least one roundup has already recirculated it as a fresh administration action. Second, the denial is categorical. A Commerce spokesman told Semafor the department “has never once asked Kalshi to take down this market or any other markets” and that “this story is false.” Kalshi declined to comment; the CFTC did not respond. So what is established is that a product existed, that it is gone, and that a news organisation and a federal department give incompatible accounts of why.
Why it ranks anyway: every compute figure this brief has carried this month — Microsoft's 38 GW by 2032, Anthropic's $13.7B six-year deal, the SemiAnalysis TCO work in item 05 — is an estimate built on private rental prices, because there is no public forward curve for GPU-hours. A traded curve would be the first continuously updating public price for the input that the whole industry's capital plans are denominated in. Whether that is a thing governments will tolerate is now a live question rather than a hypothetical one, and it is the reason to watch the CFTC docket rather than the denial.
Sources Semafor (original, with the Commerce denial) · Techmeme (summary) · Gizmodo · Crypto Briefing
Google's ATLAS: half of surveyed scientists use AI daily, and the bottleneck moved to the lab bench
Published 15 September. ATLAS is Google's standing project on AI adoption, released as an open, interactive dataset. The occupational cuts are the part most people will quote — arts, design and media make up 19% of work-related AI use in India, 1.6× the global average, while computer and mathematical occupations are 30% of U.S. use, double the global share; Brazil and the UAE adopt faster than GDP per capita predicts. Useful, but the science component is the one that should change your picture.
The science study and what it is. Run with Google DeepMind and MIT FutureTech: a survey of 600+ scientists plus an analysis of 2,600 specialised AI models. Findings as stated: nearly half of surveyed scientists use AI daily; both general LLMs and domain-specific models are in widespread scientific use; respondents report saving almost seven hours a week. Note the shape of that last number — it is self-reported time-saving from a survey, not a measured throughput gain, and the well-documented gap between the two is exactly what METR's developer-productivity work has spent the past year on. Treat it as what scientists believe about their own workflow.
The finding with teeth. The study identifies the emerging constraint as physical experimentation and clinical validation — a growing backlog of hypotheses generated faster than they can be tested. If that holds, it is the same shape as a pattern this brief keeps circling from other angles: the cheap, closed, verifiable part of an intellectual task saturates, and the expensive, open, physically-gated part becomes the rate limiter. FrontierMath Tier 4 against FrontierMath Erdős was that in mathematics; this is the wet-lab version. That connection is this brief's inference, not the study's claim.
The conflict to hold in mind. Google is both the publisher and the vendor whose product adoption is being measured, and the LLM-usage data comes from Gemini logs. The specialised-model analysis is a separate corpus. This is not a reason to discount the study; it is a reason to read the methodology note before citing a share-of-usage figure as a market statistic.
Sources Google: AI & Economy ATLAS (primary, 15 Sept) · GCN on the 15M-interaction analysis · @emollick (the jagged-frontier diagram) · METR research (on self-reported vs. measured productivity)
SemiAnalysis put its own harness on pre-release Rubin and got 67× performance per dollar
Published 14 September — outside this brief's window and missed in edition 004, which is why it runs here rather than being dropped. SemiAnalysis benchmarked Nvidia's Vera Rubin NVL72 against GB200/GB300 on its own AgentX agentic-inference benchmark, replaying real agentic traffic across thousands of chips, and reports 67× better performance per dollar versus GB300 in the agentic scenario, ~7× token throughput per megawatt against Nvidia's own claimed 3×, roughly 2× annual profit per gigawatt, and about 10× more tokens per dollar than B300 at 80 tokens/second under an ownership TCO model.
What is being measured, and against what. The workload is DeepSeek V4 Pro (1.6T parameters) under agentic replay, scored on throughput, P90 time-to-first-token, end-to-end latency and cost. The dollar figures run through SemiAnalysis's own AI Cloud TCO model — capex, opex, power, financing — under both hyperscaler-ownership and cloud-rental assumptions. That model is proprietary and its assumptions are not independently checkable, so the 67× is an independent measurement fed into a private cost model, which is a different epistemic object from a raw throughput number. The throughput and latency figures are the more portable ones.
The caveat SemiAnalysis states itself. The runs used early pre-release Rubin software. Kernel and scheduler maturity typically moves inference numbers materially in both directions over a launch cycle. SemiAnalysis's own framing is that Nvidia has historically under-claimed at GTC; that is their characterisation of Nvidia's marketing, not a fact this brief can verify.
Why the number is not as absurd as it reads. 67× is performance-per-dollar in one agentic scenario against a specific prior part, not a generational speedup. The headline efficiency claim in Nvidia's own materials is framed as up to 30× more work per watt. Any comparison between the two figures has to fix the workload, the precision and the TCO assumptions first; the versions and baselines differ.
Sources SemiAnalysis: Rubin NVL72 Agentic Inference · SemiAnalysis: Vera Rubin vs GB200 TCO · Nvidia's own efficiency claims · ServeTheHome (Hot Chips 2026 hardware detail)
Also on the wire
Confirmed, but not enough on its own to change the picture.
-
Mistral and Mozilla are putting a French frontier model inside Firefox (16 Sept)
Announced this morning: Firefox Smart Window, in beta, powered by Mistral models — summarising complex searches, recovering things you clicked away from, and organising across tabs. Launching first in France and North America, with the UK and Germany later in 2026. The terms are the interesting part: conversations are not saved to Mozilla's servers by default and Mistral commits to zero data retention, with models fine-tuned on regional languages and dialects. Read it against edition 004's item 04 — the same retention question that made Palantir, Nvidia and Booz Allen restrict frontier models is now a consumer marketing position.
Source Mistral (primary) · @MistralAI
-
Gemini now runs Siri, and it shipped on 14 September — not the 15th
iOS 27 went out on 14 September with the rebuilt Siri in English beta on iPhone 15 Pro and newer, running Apple foundation models developed with Google on Gemini, split between on-device and Private Cloud Compute, with a system orchestrator driving the Spotlight index and cross-app actions. French, Japanese, Korean, Portuguese and Spanish are slated for October; server-side features carry daily caps with paid expansion signalled. At least one roundup dated it the 15th. Whatever else it is, it is the largest single consumer deployment channel any frontier model has been handed, and the brief has been treating Apple as out of scope — that is harder to justify now.
Source TradingKey (rollout detail) · Macworld
-
Shanghai AI Lab's Atria Dawn Preview: 744B agentic MoE, weights out 11 Sept, report out 14 Sept
Open weights under MIT, built on a GLM‑5.2 base, 256K context, served via SGLang v0.5.13 or vLLM v0.23.0. The technical report (arXiv 2609.15818, v1 submitted 14 September, 143 authors) claims competitive results across 16 benchmarks and top scores on five — reported figures include DeepSearchQA 96.0, BrowseComp 92.5, CyberGym 86.5, MLE‑bench Lite 86.2, Terminal‑Bench 78.3, all self-reported. The part worth reading is the human study rather than the table: 769 task records from 56 participants with agent logs, in which roughly one-third of completed AI-assisted tasks were rated by the participant as “infeasible without AI,” with agents proposing methods and implementing revisions while humans kept most final decisions. See Checked and spiked for the dating.
Source arXiv 2609.15818 · GitHub · AI Weekly (specs)
-
Nothing from the eval operators in the window
ARC Prize's last post remains the 3 September Astra analysis; Epoch AI's last is the 14 September near-daily-use data insight carried in edition 004; METR has published no new evaluation report. Anthropic's newsroom is unchanged since 10 September and OpenAI's since 11 September. On a day when four labs' principals were arguing in public, none of them shipped a safety artefact.
Source ARC Prize · Epoch AI · METR evaluations · Anthropic · OpenAI
Checked and spiked
Items that circulated but did not survive verification.
“Anthropic cut Claude Code limits by 17%.” Carried as a 15 September item. The change is real and it took effect on 14 September, but the bare percentage does not survive contact with its baseline. Anthropic announced it on 29 August as a +25% increase — measured against the pre-promotion standard. The −17% is measured against the temporary +50% boost that ran from 13 May to 13 September and has now ended. On an indexed scale: 100 before, 150 during the promotion, 125 now. Both numbers are arithmetically correct and neither is Anthropic's characterisation of the other. And all of them are proportional by necessity: Anthropic publishes no absolute token figures for these limits, so nobody outside the company can say what any of the three levels is in units of work. If you are reasoning from this to anything about Anthropic's compute position — and several people are — note that the only stated reason on the record is “what we can sustainably serve going forward,” and that this brief has already been corrected once for inferring a capacity story out of a pricing move.
Sources Digital Applied (indexed breakdown) · Notebookcheck (the 29 August announcement)
“Atria Dawn, a 15 September model release.” It ran in at least one daily digest under yesterday's date. The model is real and worth attention (see Also on the wire), but nothing about it happened on 15 September. The weights went up on Hugging Face on 11 September; the technical report was posted to arXiv as v1 on 14 September; the only 15 September event is the author's submission of the paper to the Hugging Face papers index, which is a listing action, not a release. A tracker timestamp is not a launch date, and this is the third time in five editions that a recirculated item has arrived with the wrong one attached.
Sources arXiv 2609.15818 (v1, 14 Sept) · Hugging Face paper page (submitted 15 Sept) · LLM Stats release ledger (11 Sept)
Corrections
Errors in this brief — fixed in place above, logged here.
No corrections were issued for edition 004, and none has been flagged against edition 005 as of compile time. The correction to edition 003's thirty-hour framing is archived with edition 004 below, and the two logged in edition 002 — the ChatGPT Pro margin claim and edition 001's GPT‑6 Astra launch date — are archived with that edition. All amended passages remain marked in place.