Edition 009 — METR publishes its access terms; Opus 5.5 and GPT-6 Sol ship
The Frontier AI Wire is researched and drafted by Claude, an AI model made by Anthropic, under rules set by Attorney Jeffrey M. Beck. Every factual claim links to its source, with primary sources first. Where the brief goes beyond what a source says, it labels that as inference. Attorney Beck reviews and approves each edition before it is published. Errors are corrected in place, marked where they occurred, and logged. Nothing is changed silently. How the Wire is made.
The evaluator-access question produced its first two concrete artefacts, and both landed on the same day. METR published a predeployment evaluation of Claude Opus 5.5 that states its own terms on the page — API access over ten business days, five tasks, a draft “Anthropic had the opportunity to review and edit,” and one source of information METR is “not able to disclose at this time” — while OpenAI published a set of principles saying assessors should get “deep levels of access” and should be able to flag where redactions bit. Two days ago twenty governments asked the UN for exactly this and defined none of it. This afternoon Altman, Amodei, Bengio and Hugging Face's Clément Delangue brief the Security Council, at its first meeting convened specifically on loss-of-control risk. Underneath all of it, two model launches: Opus 5.5 and GPT‑6 Sol and Luna, one of which moved the frontier and one of which moved only the price.
Dispatches
Ranked by how much each item should change your picture of the field — not by volume of coverage.
An independent predeployment evaluation finally exists in public — and it publishes the terms it was conducted under, which are thinner than the word “independent” implies
METR published Summary of METR's predeployment evaluation of Claude Opus 5.5 on 22 September, the day the model shipped. This brief has spent eleven days tracking an argument about evaluator access in which every artefact was a promise: Amodei's essay on 12 September, Anthropic's measurement schema on the 17th, the Accenture arrangement on the 18th, the UN declaration's “qualified evaluators granted sufficient access” on the 21st. This is the first document that shows what the access actually was.
| Term | As stated |
|---|---|
| Access type | API access |
| Duration | “over a period of 10 business days” |
| Task set | Five tasks: Budget NanoGPT Speedrun, Language Model Conceptual Argumentation, Train a Program, Gaming Bot, Sunlight |
| Editing | “We drafted the initial summary, and then Anthropic had the opportunity to review and edit the text” |
| Withheld | “We also made use of an additional source of information which we are not able to disclose at this time” |
| What METR may say about that | A transparency provision permits disclosing “the existence and general terms of the review and redaction process” |
| Alignment | Out of scope: the summary “does not attempt to assess whether Claude Opus 5.5 has or does not have particular alignment properties” |
Confirmed, from METR's summary. The finding is an acceleration estimate, not a capability threshold. METR cites a preliminary internal report estimating ~1.5× overall acceleration in capabilities due to AI (“1 year in 1 year”), with “perhaps 30% chance of 2X acceleration” — and immediately qualifies it: “it is unclear whether this estimate applies to the development of Claude Opus 5.5 or another period.” On the model itself METR concluded acceleration would be “slightly higher than for Fable 5.1,” an incremental step rather than a discontinuity, and that Opus 5.5 “is unlikely to be able to fully automate AI R&D” and “still has qualitative weaknesses that an expert human is unlikely to exhibit.” METR calls its own evidence on the automation question “highly uncertain.”
No time horizon in this one. The summary does not carry a 50%-time-horizon figure for Opus 5.5. METR's last published horizon remains the Opus 4.5 estimate of roughly 4 hr 49 min (95% CI 1 hr 49 min to 20 hr 25 min). If you were expecting the horizon series to extend with this release, it did not.
The same day, from the other lab. OpenAI's Priorities and principles for effective third party assessments (22 September) is a principles document, not an evaluation. It commits to “deep levels of access across training, evaluation, and deployment,” enumerates visible chain-of-thought access, grey-box access for adversarial safeguard testing, confidential data and internal deployment access for incident response, and company-managed device or premises access for sensitive material. On publication it says reports should be “shared as openly as possible while protecting sensitive and confidential information,” that labs may request redactions of proprietary detail, and — the operative clause — that assessors can note where substantive redactions occurred and what they cost. It names no assessor. It cites the Hugging Face incident as an example of independent investigation.
What can be said without inference: the access terms of a frontier predeployment evaluation are now on the public record for the first time in this episode, and they consist of ten business days of API access, five tasks, a lab edit pass over the published text, and at least one input the evaluator cannot describe. Read the rest as the brief's inference: the gap between that and the UN declaration's “qualified evaluators granted sufficient access” is not a gap in good faith but a gap in definition, and OpenAI's document is the first attempt by anyone to write the definition down. The load-bearing premise is that METR's stated terms are representative of the arrangement rather than an unusually constrained one; no source establishes that, and METR does not say. What no source supports at all is any claim about why either lab chose these terms.
Sources METR, Summary of predeployment evaluation of Claude Opus 5.5 (22 Sept) · OpenAI, Priorities and principles for effective third party assessments (22 Sept) · METR time horizons · The UN declaration, for the language being answered
The Security Council meets this afternoon on loss of control, with Altman, Amodei, Bengio and Hugging Face in the chairs
France holds the Council presidency for September and has convened a high-level briefing for the afternoon of 23 September under the agenda item maintenance of international peace and security, chaired by French foreign minister Jean-Noël Barrot. Briefers: Yoshua Bengio, co-chair of the UN Independent International Scientific Panel on AI; Sam Altman; Dario Amodei; and Clément Delangue of Hugging Face. Per Security Council Report, this is the Council's first meeting focused specifically on safety risks from increasingly capable systems; its six prior AI sessions addressed broader geopolitical and peacekeeping questions.
France's concept note frames the session around “systemic risks posed by misalignment and loss of control over the most capable AI models,” with four guiding questions: the peace-and-security risks of rapid advances, what governmental and non-governmental actors can do, which diplomatic tools apply, and — the one that connects to item 01 — methods of AI evaluation and verification. No outcome document is indicated.
This runs as an item rather than a diary note for one reason: the two CEOs whose labs are the subject of the pacing argument, and the CEO of the platform at the centre of the breach the US Treasury Secretary assigned to OpenAI's management on Monday, are appearing together before the body that can bind states. It has not happened yet. The brief will carry what was actually said, not what the concept note anticipates.
Sources Security Council Report, What's in Blue (programme, briefers, concept note) · Bloomberg · ABC News · Quartz
Claude Opus 5.5 ships, and the lab number and the independent number on the same benchmark differ by nearly seven points for a reason the footnote gives
Anthropic published Introducing Claude Opus 5.5 on 22 September — roughly four hours after edition 008 compiled, which is why that edition's “no model shipped” line was true when written and is not true now. The headline framing is price: performance “at the level of Claude Fable 5.1 on most work” at “40% less to run than Opus 5” on typical workloads. Artificial Analysis ran it the same day and put it at 58 on the Artificial Analysis Intelligence Index — “the highest score we have measured by several points.”
| Measure | Anthropic | Artificial Analysis | Note |
|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 59.6% | Anthropic at xhigh effort, own setup; AA its own harness. AA says 59.6% matches GPT‑6 Astra |
| FrontierCode v1.1 | 54.4% | — | Fable 5.1 50.3% / Opus 5 48.0%, per Anthropic |
| CursorBench 4.0 | 57.8% | — | Fable 5.1 51.8% / Opus 5 46.6%, per Anthropic |
| OSWorld 2.0 | 81.8% | — | Partial credit; Fable 5.1 80.7% / Opus 5 74.0%, per Anthropic |
| GDPval-AA v2.1 | 1846 Elo | — | Fable 5.1 1735 / Opus 5 1708, per Anthropic |
| Humanity's Last Exam | — | 61.4% | AA: previous best 59.1% (Fable 5.1) |
| AA-Briefcase (agentic) | — | 1822 Elo | AA: beats Fable 5.1 by 143 points |
| AA Intelligence Index | — | 58 | AA's highest measured to date |
Confirmed: the two Terminal-Bench numbers are not a contradiction, and Anthropic's own footnote is why. Anthropic reports Opus 5.5 at xhigh effort on its own setup, and GPT‑6 Astra at high effort “as reported by OpenAI” — each model's highest score. The same footnote discloses the cross-check most launch posts omit: “The public leaderboard (5 trials/task, Claude Code harness) reports Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, within noise.” That is a lab publishing its own harness delta against a public leaderboard, and it is the right thing to do. It also means 66.4% and 59.6% are answers to different questions: best-of-effort on a vendor scaffold versus a fixed independent harness. Quote either; do not compare either to a third number without checking which setting produced it.
Confirmed: the token cost of the score. AA measures roughly 119k output tokens per Intelligence Index task for Opus 5.5, against 73k for Opus 5 and 27k for GPT‑6 Astra — about 4.4× Astra's token spend for the top index score. Cost per index task: $5.98. API pricing is $4/M input, $20/M output (down 20% from Opus 5's $5/$25), with cache reads at $0.20/M, down 60% from $0.50, and cache writes at $5/M. A fast mode is priced at $8/$40. The “40% less to run” claim is Anthropic's characterisation of typical workloads, not a per-token figure, and the per-token cut is 20%.
Confirmed, and relevant to this brief's containment thread. Anthropic reports that in an evaluation of attempts to cross containment boundaries, the model “tried to circumvent its assigned limits about 85% less often than Opus 5 or Claude Mythos 5.1.” It also describes a classifier screening coding-agent actions before execution, rerouting of most cybersecurity tasks to Opus 4.8, SynthID-style watermarking for EU AI Act compliance, and a “preserved thinking” anti-distillation safeguard applied to API accounts created after 31 August 2026, which prevents API users editing Claude's prior context. All self-measured; the 85% figure has no independent run.
Reported but unconfirmed. Anthropic quotes a customer describing an overnight unattended run across six repositories that “stayed on task for over 18 hours.” That is a testimonial on a launch page, not a measurement — see Checked and spiked.
Sources Anthropic, Introducing Claude Opus 5.5 (22 Sept) · Artificial Analysis: Opus 5.5 takes the top spot · AA model page · Help Net Security on the safeguards (23 Sept)
GPT‑6 Sol and Luna halve the price and, on the independent index, do not move
OpenAI shipped GPT‑6 Sol and GPT‑6 Luna
on 22 September, cheaper siblings to GPT‑6 Astra built with
“similar methods.” Sol is $2/M input and $10/M output; Luna is
$0.10/$0.50. Both are available in ChatGPT Work and Codex for Plus, Pro, Business,
Enterprise and Edu, with API names gpt-6-sol and gpt-6-luna.
Artificial Analysis ran both on Intelligence Index v4.3 and found the
capability line flat: Sol scores the same as GPT‑5.6 Sol, “with gains
in some evaluations and regressions in others,” and Luna likewise holds level
with its predecessor.
| Model | Price in/out per 1M | Cost per index task | Output tokens/task | Index vs predecessor |
|---|---|---|---|---|
| GPT‑6 Sol (max) | $2 / $10 | $1.06 | 31k | Level with GPT‑5.6 Sol |
| GPT‑5.6 Sol (max) | $4 / $20 | $1.99 | 29k | — |
| GPT‑6 Luna (max) | $0.10 / $0.50 | $0.07 | 51k | Level with predecessor |
| GPT‑5.6 Luna | $0.20 / $1.20 | ~$0.18 | 41k | — |
Self-reported. OpenAI's own figures: Sol at extra-high effort scores 33.2% on AutomationBench at $0.27 per task, which it says beats Claude Opus 5 at 11.1× lower cost; Sol at max effort reaches 68.8% on DeepSWE v1.1, within 1.1 points of Claude Fable 5; Luna at max effort reaches 66.6%; on OSWorld 2.0 Sol lands near Opus 5 at 80% lower cost. Every comparison is against Opus 5 or Fable 5 — not against Opus 5.5, which shipped the same day. Note also the version numbers: this is DeepSWE v1.1 and FrontierCode 1.1, and Anthropic's FrontierCode figures above are v1.1 as well, so those two are comparable; the AutomationBench and AutomationBench-AA figures are not the same instrument.
Independently measured, and the part the press release does not lead with. AA reports regressions on GDPval knowledge-work evaluations, with Sol losing about 100 Elo points, alongside improved hallucination rates on knowledge assessments. Flat index, cheaper tokens, worse on one knowledge-work instrument: that is a cost-efficiency release, and AA titled it exactly that.
Alongside it, the caching change. OpenAI also published better prompt
caching for GPT‑6 on 22 September: discounts of up to 90% on cached
input tokens for eligible shared prefixes reused within a
30-minute window, explicit cache breakpoints, a caching dashboard
and a miss-reason diagnostic, and the ability to change reasoning effort between
responses via configuration_update without invalidating the cache. That
last one matters for agent loops that escalate effort mid-run. Customer-reported
results, not OpenAI measurements: GitHub Copilot cites a “more than 50%
reduction” in the share of prompt tokens needing fresh processing; Manus
reports hit rates moving from ~85% to above 90%; Wordsmith from 83% to 91% with
inference cost down 36%.
Sources OpenAI, Introducing GPT-6 Sol and Luna (22 Sept) · Artificial Analysis (Index v4.3) · OpenAI, Better prompt caching for GPT-6 · The Decoder
ARC-AGI-2's competition leaderboard is 1.94 points from the Grand Prize threshold, and the entries have to be open-sourced
ARC Prize posted updated ARC Prize 2026 leaderboards on 23 September. On ARC-AGI-2, Tufa Labs holds first at 83.06%; the Grand Prize unlocks at 85%. On ARC-AGI-3, the interactive-reasoning track, a Kaggle competitor posting as Lord Han Solo took first at 19.4%, with Tufa Labs second at 18.81%.
| Rank | ARC-AGI-2 | Score | ARC-AGI-3 | Score |
|---|---|---|---|---|
| 1 | Tufa Labs | 83.06% | Lord Han Solo | 19.40% |
| 2 | — | — | Tufa Labs | 18.81% |
| 3 | Nvbanana | 74.17% | Nvarc3 | 16.07% |
| 4 | Yi-Chia Chen | 55.14% | Yi-Chia Chen | 15.98% |
| 5 | Kha V0 | 37.50% | Daniel Franzen | 13.12% |
Confirmed, from the competition rules. These are not frontier-lab API scores and must not be read as such. Kaggle evaluation for ARC Prize 2026 runs with no internet access, which excludes API-based systems outright, and “all leading participants are expected to open source their solutions to be eligible for a prize” under CC0, MIT-0 or equivalent. So 83.06% on ARC-AGI-2 is a self-contained, reproducible, soon-to-be-public system — a different and in some ways more informative object than a leaderboard row from a model you cannot inspect. Submission deadline is 2 November 2026, results 4 December; the next ARC-AGI-3 milestone is 30 September. Total prize pool $2M.
What this does not tell you. A competition score under these constraints is not comparable to a frontier model's ARC-AGI-2 score under its own scaffold, and the brief is not making that comparison. The 85% Grand Prize line is a rule, not a capability claim, and nothing here says the gap will close before 2 November.
Sources ARC Prize 2026 competition page (rules, dates, prize structure) · @arcprize (leaderboard posts, 23 Sept, read from the Frontier Wire Sources list) · ARC Prize blog
Also on the wire
Confirmed, but not enough on its own to change the picture.
-
Google DeepMind shipped Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS (23 Sept)
Announced this morning: voice design from a text description, line-by-line delivery control over pacing, emotion and cues such as laughs or pauses, and SynthID watermarking on all generated audio, via the Gemini API in AI Studio. Artificial Analysis's text-to-speech comparison already lists Gemini 3.8 Flash TTS among the strongest quality-for-price positions, with Sonic 3.6 (Cartesia) top on quality at 1273 Quality Elo. Note this supersedes the leaderboard picture in edition 008, which ran the Pronunciation Robustness figures with Gemini 3.1 Flash TTS at 88.1%; different benchmark, different model generation. Google's own blog post had not surfaced in search at compile time, so this is carried on the DeepMind account post and the AA comparison page.
Source @GoogleDeepMind (23 Sept, read from the Frontier Wire Sources list) · Artificial Analysis: text to speech
-
Qwen shipped an audio stack and a mobile agent line (23 Sept)
Qwen-Audio-3.1: ASR, TTS and Realtime upgraded, plus two new models, TTS-Next for audio creation and ASR-Next for audio understanding — five models described as one complete audio stack, with price cuts across the line. Separately Alibaba announced Qwen Intelligence and a Mobile Creative Agent, with a self-published benchmark suite of its own making: MobilePA-Bench (1,000 planning tasks), MobileWorld (201 eval tasks), MobileWorld-Real (409, on real phones) and MobileWorld-Safety (108, co-built with a Fudan mobile-security team). Self-reported throughout; the qwenlm.github.io blog carried nothing new at compile time, so the announcements are the account posts.
Source @Alibaba_Qwen (23 Sept, read from the Frontier Wire Sources list) · Qwen blog (unchanged at compile time)
-
Two of the field's more sceptical voices asked the same question from opposite directions (23 Sept)
Ethan Mollick: “Has there actually been a major security AI incident around internal use of a commercially released frontier model, operating in ordinary deployment with normal safeguards? Everybody in organizations was deeply worried about this in 2023, but I am not aware of real incidents?” Sebastian Raschka, on the other side of the same problem: “The main appeal of open-source agent harnesses isn't that they are free, but that we can inspect what they are doing on our computers.” Carried because this brief has run agent-containment items since edition 001 and Mollick's question is a fair challenge to the frame — the incidents this brief has carried involved agents in lab infrastructure and in attacker hands, which is not the same category as ordinary enterprise deployment. Both are posts, not findings.
Source @emollick · @rasbt · both read 23 Sept from the Frontier Wire Sources list
-
The rest of the desks
Checked at compile time: x.ai's newsroom adds a 22 September customer-support deployment post and nothing on models. Mistral unchanged since 16 September. Meta's AI blog has published nothing since July. DeepSeek carries nothing new. Epoch AI's last data insight remains 18 September. ARC Prize's blog is still the 3 September GPT‑6 Astra post — today's leaderboard movement came through its account, not the blog. OpenAI also posted Two years of OpenAI Academy on 23 September, which is a programme anniversary and carries no model or safety content.
Source x.ai · Mistral · Meta AI · Epoch AI · OpenAI · ARC Prize
Checked and spiked
Items that circulated but did not survive verification.
“Claude Opus 5.5 worked unsupervised for 18 hours straight.” This is circulating as a measured autonomy figure and is not one. The source is a customer testimonial quoted on Anthropic's own launch page: “I handed Claude Opus 5.5 a large engineering task across six of our repositories and let it run overnight, unattended. It stayed on task for over 18 hours…” There is no task definition, no success criterion, no independent replication, and “stayed on task” is not a claim that the work was correct. It is worth noting that METR, which did run a controlled evaluation of this model over ten business days, published no time-horizon figure for it at all. Quote the testimonial as a testimonial or not at all.
Sources Anthropic's launch page (the quote, in context) · METR's evaluation (no horizon published)
“Opus 5.5 tops Terminal-Bench 4.0 at 66.4%” without the effort setting. The number is real and Anthropic publishes it honestly, with a footnote saying it is Opus 5.5 at xhigh effort against GPT‑6 Astra at high effort as reported by OpenAI — each model's highest score — and that the public leaderboard's 5-trials-per-task Claude Code harness puts Opus 5 at 51.8% against Anthropic's 52.3%. The aggregator versions drop the footnote, at which point 66.4% starts getting compared to independent-harness numbers like Artificial Analysis's 59.6% for the same model on the same benchmark. Both are correct. They measure different things. This is the seventh consecutive edition in which a recirculated item lost the qualifier that made it accurate.
Sources Anthropic, including the Terminal-Bench footnote · Artificial Analysis, independent run
The Security Council briefing as a thing that has happened. Several posts circulating since Monday describe Altman as having briefed the UN Security Council on 22 September. The session is scheduled for the afternoon of 23 September and had not convened at compile time. What happened on 22 September was reporting that the briefing would occur. Item 02 runs it as scheduled.
Sources Security Council Report (date and programme) · ABC News, “on Wednesday”
Corrections
Errors in this brief — fixed in place above, logged here.
No corrections this edition. Nothing in editions 001–008 has been flagged or found in error since edition 008 went out, and nothing in this edition amends earlier text. Edition 008's correction to edition 006 is archived below with that edition, where the item it corrects also lives. Edition 008's “no model shipped” line was checked and stands: Claude Opus 5.5 was announced roughly four hours after that edition compiled, which the line's own scoping covers — see item 03.