Edition 018 — OpenAI Publishes the Limits of Its Own Watermark

The Frontier AI Wire is researched and drafted by Claude, an AI model made by Anthropic, under rules set by Attorney Jeffrey M. Beck. Every factual claim links to its source, with primary sources first. Where the brief goes beyond what a source says, it labels that as inference. Attorney Beck reviews and approves each edition before it is published. Errors are corrected in place, marked where they occurred, and logged. Nothing is changed silently. How the Wire is made.

A frontier lab shipped text watermarking and published the numbers that bound it. OpenAI's textGrain went live yesterday as a global opt-in for API customers and reaches ChatGPT and Codex output in the EU “over the coming weeks” — and the post states, unprompted, that at a 1% false-positive rate the detector finds the mark in about 95% of 400-token passages, that replacing 10% of words with synonyms takes that to 66%, and that 25% takes it to 17%. Separately, Reflection announced Beam — 501B parameters, 23B active, Apache 2.0 — and did not release it; the weights come “later this month,” and on Reflection's own table Beam trails Kimi K3 and DeepSeek V4.1 on the agentic benchmarks it chose while claiming 3–4× less inference compute than GLM 5.2. The Information reported two of the largest American technology companies cutting internal Claude use. OpenAI's first-party web search entered Artificial Analysis's search board fifth. And forty-eight hours on, a presidential action creating the Super Intelligence Force was still not found at compile time. No new corrections; three items circulating this week are months old or unsupported.

Dispatches

Ranked by how much each item should change your picture of the field — not by volume of coverage.

01
Primary source Shipping, not announced Published measurements No independent run

OpenAI puts a text watermark into production for the EU — and publishes the detection rates that bound what it can prove

Posted to OpenAI's newsroom on 5 October, under the title Our approach to EU text provenance rules. The mechanism is called textGrain, and OpenAI describes it in one sentence: it “adds an invisible statistical signal to the model's word choices.” The stated driver is the EU AI Act, which OpenAI says “requires generative AI providers to make generated text identifiable in a machine-readable way.” Three rollout steps, in the post's own order: “Starting today, API customers globally will be able to opt in to text watermarking” for selected models; “over the coming weeks, we will add an invisible watermark to eligible ChatGPT and Codex text output in the European Union”; and applications open for the detector, with access “initially limited to approved researchers.” OpenAI says it is C2PA conformant and “plans to make the technology available in open source.”

Detection, as OpenAI reports it — all at a 1% target false-positive rate
ConditionWatermark foundNote
200-token passage~80%The shorter of the two lengths the post reports
400-token passage~95%The headline figure
400-token passage, 10% of words replaced with synonyms92% → 66%OpenAI's own wording. Note the 92% baseline for this experiment, which is not the 95% above
400-token passage, 25% of words replaced17%Same experiment
Independent evaluationnoneNot found at compile time at artificialanalysis.ai or metr.org/blog
Technical detail — worth digging further

The brittleness number is the finding, and the vendor published it. A statistical watermark in word choice survives being read and does not survive being edited: a tenth of the words swapped for synonyms takes detection from about 92% to 66%, and a quarter takes it to 17%. Every public argument about AI-text detection for the last two years has run without a number attached. There is now a number, it comes from the party shipping the mechanism, and it says the mechanism is a provenance signal for unmodified output rather than a test anyone can apply to a document of unknown history.

What OpenAI says a watermark cannot do, verbatim, because the post is unusually explicit about it. “A watermark does not measure human contribution”; “a watermark does not establish ownership or responsibility”; “a watermark does not identify the user”; “a watermark does not verify accuracy”; and “the absence of a detected watermark does not prove human authorship.” Five disclaimers in a launch post is not the usual ratio, and the last one is the one that governs how the detector can be used: a negative result is uninformative.

The access design is the second thing to watch. The watermark is going on by default in one jurisdiction and by opt-in everywhere else, while the detector is gated behind an application process restricted at first to approved researchers. That is an asymmetry between who gets marked and who can check, and it is the shape every evaluator-access argument in this brief has taken since edition 009 — a capability exists, and the terms on which outsiders may use it are set by the party that built it. OpenAI's stated intention to open-source the technology would change that; nothing published says when, and the post does not say whether open-sourcing covers the detector, the embedder, or both. OpenAI notes separately that its verification tools for audio and images, including openai.com/verify, remain publicly accessible.

What this brief has not established. No independent evaluation of textGrain was found at compile time at either of the two eval operators that would be likeliest to run one. Nothing here establishes the false-negative rate on text a model produced and a person then rewrote in their own words rather than by synonym substitution, which is the realistic case and is not among the conditions reported. And on the one comparison a reader will reach for: edition 009 recorded Anthropic describing “SynthID-style watermarking for EU AI Act compliance” shipping with Claude Opus 5.5. That description does not say whether it covers text, and this brief has not established that it does, so the two should not yet be set against each other.

Sources OpenAI, Our approach to EU text provenance rules (primary, 5 Oct) · OpenAI newsroom (dating; nothing dated 6 Oct at compile time) · ActuIA, on the opt-in asymmetry (independent dating) · AI Weekly (independent dating)

02
Primary source Unusual training disclosure Weights not released Self-reported benchmarks

Reflection announces Beam, a 501B open-weight model it has not shipped — and the claim is efficiency, not capability

Announced 5 October. A sparse mixture-of-experts model, 501B total parameters with 23B active, 52 layers, fine-grained routed experts with interleaved local and global attention, and a 1,000,000-token context extended during midtraining. Pre-trained on 23.8 trillion tokens. The licence is promised, not granted: Reflection's own status line reads “Beam is undergoing final red-teaming and evaluations… We will release the weights, technical report, model card, and developer artifacts later this month.” Nothing is downloadable today.

Beam against the open frontier, as Reflection's own post states it
BenchmarkBeamComparators, as Reflection reports them
DeepSWE v1.144.4GLM 5.2 44.0 · Qwen 3.8 Max 51.0 · Kimi K3 68.0 · DeepSeek V4.1 74.2
Terminal‑Bench v2.180.1GLM 5.2 81.0 · Qwen 3.8 Max 86.6 · Kimi K3 88.3 · DeepSeek V4.1 90.6
SWE‑Bench Pro v2‑Hard77.2GLM 5.3 84.3 · Kimi K3 88.2
SWE‑Bench Pro v165.5Qwen 3.8 Max 67.7
SWE‑Bench Verified80.9Inkling 77.6 · Nemotron 3 Ultra 70.7
Independent evaluationnoneNo Beam entry found at compile time at artificialanalysis.ai; TechCrunch: “Reflection's performance claims haven't been independently verified”
Technical detail — worth digging further

Read the table before the headline, and note who published it. On Reflection's own numbers Beam is roughly level with GLM 5.2 and behind Kimi K3 and DeepSeek V4.1 on the agentic benchmarks Reflection selected. The post says so itself: “Where frontier open models like Kimi K3 remain ahead on raw capability, Beam's advantage is efficiency at inference time.” A lab publishing a comparison table it loses on is rare enough to be worth naming as a fact about the document rather than about the model.

The efficiency claim, and how it is computed. “Beam achieves scores comparable to GLM-5.2 while using 3–4× less inference compute.” The proxy behind that is stated: “FLOPs ≈ 2 × active parameter count × mean generated tokens per attempt,” with the exclusions also stated — “these estimates exclude prompt prefill, context-dependent attention operations, and serving overhead.” Two things follow. First, a metric that is a product of active parameters and token count rewards a low active-parameter count by construction, which is what a 23B-active model has; that is an observation about the instrument, not an allegation about the result. Second, excluding prefill and context-dependent attention is a larger omission for a model whose headline feature is a million-token context than for one without. Read the 3–4× as a statement about arithmetic per generated token, not about what a request costs — and the load-bearing premise for treating it as the latter is that the excluded terms are small, which Reflection does not claim and nobody has measured.

The operational disclosure is the part to keep. Pretraining ran on 6,144 NVIDIA GB300 NVL72 GPUs in under four weeks at 92.3% goodput, with nine semi-automatic checkpoint rewinds. The high-compute RL run generated “over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks,” across roughly 1.3 billion sandboxes and about 110,000 sustained concurrent rollouts, at up to 256K context, with stability maintained at one day of policy staleness — 107 weight versions behind the current policy. Reflection also publishes a median weight-distribution time of about 12 seconds across the inference fleet and a median inference-failure recovery of 8 minutes, with 0.02% of serving GPU-minutes lost. Set against the comparators: Reflection puts Inkling at 30M rollouts and MiMo at 753K. No frontier lab in eighteen editions of this brief has published operational numbers at this granularity, and they are checkable in a way the benchmark table is not — anyone running a comparable job can say whether 92.3% goodput and one-day staleness tolerance are plausible.

What is not established, and it is most of it. The weights. Every performance figure is Reflection's own; TechCrunch states plainly that the claims have not been independently verified, and Artificial Analysis carried no Beam entry when checked at compile time. Until the weights land, this is a specification and an engineering write-up, and this brief does not rank unmeasured numbers above measured institutions. The financing figures in circulation are TechCrunch's reporting and are carried as such: roughly $7 billion in compute agreements for GB300 access through 2029 — over $6B with SpaceX and $1B with Nebius — and about $4.7 billion raised, most recently at a reported “$25 billion pre-money valuation.” Nothing in any of that supports a statement about Reflection's costs, runway or strategy, and none is made here.

Sources Reflection, Introducing Beam (primary, 5 Oct) · TechCrunch, 5 Oct (independent dating, financing figures, the verification caveat) · Implicator, 5 Oct · Artificial Analysis (checked at compile time — no Beam entry)

03
Reporting Figures consistent across accounts Original paywalled to this brief No company response

Microsoft and Meta are cutting internal Claude use, per The Information — a third off one budget, half the users off the other

Reported 5 October by The Information and carried onward the same day by Finimize, StockTwits via Yahoo Finance, AI Weekly and others. Two figures appear in every secondary account this brief could read, in the same words: Microsoft has “cut projected internal spending on Claude by more than a third,” and Meta's internal Claude Code user count has fallen to “about 30,000 from roughly 60,000 earlier this year.” The reasons given in the reporting are cost control, data privacy, and a preference for keeping work on internal or partner tools. The named alternatives: Microsoft directing employees to OpenAI models through GitHub Copilot; Meta deploying MetaCode and Muse Code.

Two further figures appear in one of the accounts read here and not the others, and are carried with that provenance attached: Microsoft's projected annual internal Claude spend put at roughly $1 billion before the cut, and monthly usage caps in its Cloud and AI division lowered from $100,000 to about $10,000.

Technical detail — worth digging further

What this is, narrowly. A paywalled scoop, read here only through secondary summaries that agree with one another on the two headline figures. No statement from Anthropic, Microsoft or Meta appears in any of them, and no account this brief read names a source. Agreement across syndications is evidence of faithful copying, not of the underlying reporting being right.

What no source supports, and it is the thing being written anyway. Nothing in this reporting establishes anything about Anthropic's revenue, revenue growth, customer concentration, margins or listing timetable. At least one secondary account attaches “notable headwinds” for Anthropic's anticipated IPO to these figures; to get from two customers reducing internal consumption to a statement about a company's listing you would need the share of revenue those customers represent and the direction of the rest of the book, and no account read here supplies either. See Checked and spiked. Edition 002's correction was for exactly this class of claim.

The adjacent document, and why this brief is not joining it to this one. Edition 016 carried Anthropic's leaked IPO prospectus, which states that nearly 25% of 2025 revenue came from two unnamed customers, most of whom lack long-term contracts. That is a filing's own language about concentration. It is a reason to want this week's reporting verified; it is not a basis for assuming the two unnamed customers are these two companies, which nothing published says and which is the inference this brief is declining.

What the figures do describe. Buyer behaviour, which is what the people quoted can see. If the reporting holds, two of the largest American technology companies are moving internal AI consumption onto stacks they control — which is a statement about where inference demand sits, and says nothing about what it costs anyone to serve it.

Sources Finimize, 5 Oct (the two headline figures, the named internal tools) · StockTwits via Yahoo Finance, 5 Oct (the $1bn base and the usage caps — single-source here) · AI Weekly, 5 Oct · @GaryMarcus, quoting @rohanpaul_ai — read 6 Oct from the Frontier Wire Sources list

04
Independent instrument Published methodology Open harness

OpenAI's own web search lands fifth on the independent search board — the first first-party search tool anyone has run on it

Artificial Analysis added OpenAI Web Search to its Search Index on 5 October, describing it as the first integrated first-party search tool on the board rather than a standalone search API. Note what is not new: the index itself has been running since at least 10 September, when this brief carried Octen's debut on it in edition 002 — see Checked and spiked, because the distinction is already being lost.

Artificial Analysis Search Index — leaderboard read at compile time
Provider (config as listed)Search IndexTime per task
Perplexity Search (medium)8027.5 s
Octen Search (highlights)7715.9 s
Parallel Search (advanced)7541.4 s
Brave Search (LLM context)7522.7 s
OpenAI Web Search (medium)7430.4 s
Technical detail — worth digging further

The methodology, because AA publishes it and it is what makes the number worth quoting. The index is “the equal-weighted mean of each benchmark's primary quality metric” across three evaluations: DeepSearchQA (900 tasks from Google's public eval split, scored on F1), BrowseComp (200 hard web-search samples selected from OpenAI's dataset, scored on accuracy) and AA‑Omniscience (600 private samples balanced across six domains, scored on exact-answer accuracy). Every row runs through AA's open-source Stirrup harness on a fixed agent loop with web_search and web_fetch and a 25-turn budget. Time per task is the sum of model time and search time.

The design decision worth naming out loud. The candidate answer model is held fixed at GPT‑5.6 Luna (medium) for every provider, which is what isolates the search layer from the model consuming it — the right design, and the same logic behind AA's CyberGym harness in edition 015. It also means that from this week one vendor is both the fixed answer model and an entrant on the board. No source alleges any effect from that, no methodology change accompanied the addition, and this brief is not alleging one; it is recording the fact so that anyone differencing row five against row one knows what is held constant.

Speed, and where the two figures differ. AA's thread puts OpenAI Web Search at about 31 seconds per task, 14th of 26 products, mid-pack against a board range of roughly 16 seconds (Octen) to 62 seconds (Firecrawl), and notes that OpenAI does not report search time separately, so AA measures it from the stream. The leaderboard read at compile time gives 30.4 seconds. Same measurement a few hours apart, not two different claims.

What a 74 against an 80 is and is not. Six points on an equal-weighted mean of three task sets with one fixed answer model is not a ranking of search quality in general. It is reproducible, which is the rarer property: the harness is open, the task sets are named, two of the three are public, and the configuration of every row is printed beside it. Most numbers this brief carries have none of those.

Sources Artificial Analysis Search Index (leaderboard, read at compile time) · AA's published methodology (datasets, index construction, answer model, Stirrup) · Stirrup, the open-source harness · @ArtificialAnlys — read 6 Oct from the Frontier Wire Sources list

05
Primary source Published today Named institutions No outcomes yet

A three-year programme to decide how AI for antimicrobial resistance should be tested — funded by Google DeepMind, run out of Imperial

Published 6 October by the Fleming Initiative and surfaced through Google DeepMind's repost of the announcement: the only item in this window carrying today's date. A three-year programme to establish evaluation methods and standards for AI systems applied to antimicrobial resistance — common standards defining how such systems should be tested, and what evidence demonstrates reliable performance across populations, settings and datasets. Google DeepMind provides the funding. Named: Professor Alison Holmes, Fleming Initiative director, who frames the aim as ensuring AI systems are “accurate, reliable and appropriately representative of the populations and settings where they will be used”; Agata Laydon, science lead at Google DeepMind's Impact Accelerator, on “robust, objective evaluation” as the condition for translating AI work into health outcomes; and an inaugural Google DeepMind Academic Fellow, appointed as an assistant professor at Imperial. The institutions named are Imperial College London, Imperial College NHS Trust and Google DeepMind.

Technical detail — worth digging further

Why a measurement programme in one disease area earns a dispatch. Eighteen editions have turned on a single recurring absence: a capability claim arrives, and no instrument exists to check it. In AI-for-science the absence has been close to total — edition 010 carried Anthropic's enzyme-family discovery, edition 012 the nine-loop amplitude, edition 017 thirty-six manuscripts in eighteen fields, every one self-measured by the lab that produced it and none independently replicated. This is three years and a named institution aimed at the instrument rather than at the result, in a domain where the endpoint is clinical and the populations are the thing that usually goes unmeasured.

What is not established. No published benchmark, no named evaluation, no funding figure, no list of the “priority AMR applications” the programme says it will cover, and no milestone short of the three-year horizon. The first checkable artefact is a published evaluation framework, and there is not one yet. Nothing here measures any model.

The structural caveat, and the comparison it invites. The funder is a frontier lab whose own systems are candidates for assessment under whatever standard emerges. The work is housed at a university and an NHS trust rather than at Google, and the named fellow holds an academic appointment. Set that against the two other oversight structures this brief has carried: Anthropic paying Accenture for embedded evaluators with publication rights unspecified (edition 007), and OpenAI seating nine mathematicians who “will not be paid by OpenAI” and may publish unrequested advice (edition 007). Three arrangements, three theories of what buys independence, and no published evidence on which of them produces a usable finding first.

Sources Fleming Initiative, New programme to evaluate AI systems for antimicrobial resistance (primary, 6 Oct) · Fleming Initiative news index (dating) · @FlemingCentre, reposted by @GoogleDeepMind — read 6 Oct from the Frontier Wire Sources list

06
Reporting Page checked at compile time Creating instrument still not located Outlets disagree

Forty-eight hours on, the Super Intelligence Force still has no creating instrument this brief can find — and one outlet's headline says it was an executive order

Edition 017 carried the Force as announced on Sunday 4 October by social-media post, with Director of National Intelligence Jay Clayton at its head and a report on risks, opportunities and federal responsibilities due in 120 days, and reported that a presidential action creating it was not found at compile time. Checked again this morning, and the answer has not changed. At compile time on 6 October, whitehouse.gov/presidential-actions lists two October items: a National Manufacturing Day proclamation dated 2 October and an executive order on emergency tax relief on diesel fuel dated 5 October. Neither creates the Force, and nothing else dated 1–6 October appears on that page. The renaming order of 29 September, Inaugurating the Era of Super Intelligence, is still listed and is still a different document.

Technical detail — worth digging further

The disagreement is still live, and it has sharpened. Edition 017's sources described a Truth Social announcement; at least one outlet now asserts an order in its headline — BigGo Finance, “Trump Signs Executive Order Creating Super Intelligence Force, Led by Clayton, to Deliver Risk Report Within 120 Days.” Further coverage surfaced in this window from TechCrunch and the Washington Times; this brief read their listings and did not read the pieces, and is not characterising what they say about the instrument.

What a missing instrument does and does not mean. It does not establish that no order exists. Presidential actions are sometimes posted after the fact, and this brief has no visibility into anything unpublished. What it establishes is narrower and is the point: forty-eight hours after the announcement, nobody outside the executive branch can read the document that would say what the Force is empowered to do, who sits on it, or what the clause about “preventing overregulation and regulatory capture” obliges it to do. Every description in circulation, this brief's included, remains a description of a description.

Two dated deliverables, no published charter. If the 120-day clock runs from 4 October, the report falls at the start of February 2027. The renaming order's separate obligation — sixty days for the Assistant to the President for Science and Technology to propose legislative language for a federal definition — falls at the end of November. Those are the two checkable moments in the whole SI thread.

What no source supports. That the Force has met, has staff, has a budget, or has any authority over any agency's existing enforcement; and nothing published this window moves the membership and adviser lists edition 017 flagged as unverified.

Sources whitehouse.gov presidential actions (checked at compile time, 6 Oct) · BigGo Finance (the executive-order assertion) · TechCrunch, 4 Oct (surfaced this window — not read here) · Washington Times, 4 Oct (surfaced this window — not read here) · PBS NewsHour, 4 Oct (edition 017's account of the announcement method)

Also on the wire

Confirmed, but not enough on its own to change the picture.

  • Eleven v4 Turbo takes the top quality slot on Artificial Analysis's text-to-speech board (5 Oct)

    AA's text-to-speech comparison, read at compile time, puts Eleven v4 Turbo at the highest quality position with a Quality Elo of 1334. AA's own thread adds the throughput and price: 96 characters per second of generation at $40 per million characters, against Eleven v4 at 84 characters per second and, per AA, twice the price. Recorded for the shape rather than the result — a faster, cheaper sibling taking the quality lead from the model it derives from is the inverse of the pattern this brief has now logged four times in the language-model tier, where per-token prices fall while tokens per task rise.

    Sources Artificial Analysis: text to speech (Quality Elo, read at compile time) · @ArtificialAnlys (throughput and price) — read 6 Oct from the Frontier Wire Sources list

  • Nathan Lambert on where the Western open-weight frontier sits (5–6 Oct)

    Verbatim: “Reflection joins the list of Nvidia & Thinking Machines who have released their strongest models and come up behind Chinese counterparts. There's a lot of ways people will overthink this, but the clearest takeaway should be that the Chinese are very very good at building LLMs.” A post, not a finding, and the dating matters: Thinking Machines' Inkling shipped on 15 July and Nvidia's Nemotron 3 family earlier still, so this describes a cumulative pattern rather than three releases in one window. Carried because item 02's own table is the clearest published evidence for the claim this week, and the lab that published it is the Western one.

    Sources @natolambert — read 6 Oct from the Frontier Wire Sources list · Thinking Machines, Inkling (for dating) · Nvidia, Nemotron 3 (for dating)

  • The rest of the desks, checked page by page

    Checked at compile time, each claim tied to the page it came from. OpenAI's newsroom carries two 5 October posts — the EU text-provenance post (item 01) and the advertising post that ran as edition 017's item 03 — and nothing dated 6 October. Anthropic's newsroom shows nothing after the 2 October Frontier Academy post; its research index nothing after Claude-shaped science of 1 October. Google: blog.google's AI section again returned no dated entries to this brief's fetcher, as in edition 017, and deepmind.google/discover/blog remains unordered by date with nothing datable from it, for the fifteenth consecutive edition — note that the Google DeepMind item in this edition ran on a partner's site and the lab's X account, not on either Google property. METR's blog is unchanged since Painter's 30 September testimony. ARC Prize's blog is unchanged since 3 September. Epoch AI's most recent publication remains the 2 October agent-population report carried in edition 017; its capabilities and benchmarking data pages are marked updated 5 October, and a data refresh is not a publication — see editions 014 and 017. Artificial Analysis adds the Search Index row (item 04) and the text-to-speech leaderboard change, above. Mistral nothing after 28 September; x.ai nothing after 28 September; Meta's AI blog nothing since July. Qwen's blog returned a stale listing to this brief's fetcher at compile time, its newest entry dated 23 September 2025, so nothing could be dated from it in either direction. DeepSeek's news page again could not be read directly, its content sitting behind a navigation link this fetcher did not resolve. One model was announced in this window and none shipped.

    Source OpenAI · Anthropic newsroom · Anthropic research · blog.google AI · DeepMind blog index · METR · ARC Prize · Epoch AI · Artificial Analysis · Mistral · x.ai · Meta AI · DeepSeek · Qwen

Checked and spiked

Items that circulated but did not survive verification.

“The United States has banned Chinese data-centre gear.” Circulating this week off an analysis piece dated 1–2 October. The underlying events are months old, and the piece says so. The executive order banning foreign-made transformers, high-voltage circuit breakers and other grid equipment linked to China and 23 other nations was signed on 26 August 2026 — and by that account nothing is actually prohibited until the Department of Energy identifies the risky equipment and publishes implementing rules, due 24 December 2026. The companion claim travelling with it, that Beijing now requires spouses and children of top AI researchers to obtain approval before travelling abroad, rests on Bloomberg reporting about exit and entry rules that took effect on 15 September. Both are real; neither happened in this window. The December rulemaking date is the thing to diarise, not the order.

Sources The analysis piece, with both dates in it

“Microsoft and Meta cutting Claude is a headwind for Anthropic's IPO.” The cuts are reported and run as item 03. The conclusion is not supported by anything in the reporting, and it is the same error class this brief was corrected for in edition 002. To get from two customers reducing internal consumption to a statement about a company's listing you need three things: the share of revenue those customers represent, the direction of the rest of the book, and the listing timetable. No account this brief could read supplies any of them. One secondary account attaches “notable headwinds” to the figures regardless. Edition 016's archive carries Anthropic's leaked prospectus stating that nearly 25% of 2025 revenue came from two unnamed customers — and assuming those two are these two is precisely the move this spike exists to refuse. Report the cuts; the inference is not available.

Sources The account carrying the IPO framing · A second account carrying the same figures without it

“Artificial Analysis launched a new search benchmark.” The row is new; the board is not. The Artificial Analysis Search Index has been running since at least 10 September, when this brief carried Octen Search's debut on it in edition 002 at a score of 77 — which is still its score on the leaderboard read at compile time, now with a time per task of 15.9 seconds against the roughly 17 seconds recorded then. What happened on 5 October is that OpenAI Web Search was added as the first integrated first-party search tool on it. A new entrant on an existing instrument is not a new instrument, and this is the same distinction editions 011, 014 and 017 drew about Intelligence Index configuration rows — fourth consecutive edition in which an existing board's new row is at risk of being read as a launch.

Sources The Search Index, with Octen still at 77 · The methodology page, unchanged in substance

Corrections

Errors in this brief — fixed in place above, logged here.

No new corrections. Nothing in editions 001–017 has been flagged by a reader or found in error since edition 017 went out. Editions 002, 003, 006, 008, 014 and 015 carry their own corrections, archived below with their editions — edition 015's correction, logged in edition 016, is archived with that edition. The standing note, carried forward: a statement that was true when written and overtaken by a publication hours later is not an error, and this brief will keep saying so rather than quietly retrofitting; the test is whether it is applied honestly in both directions. The rule adopted after edition 015 — that an absence is reported as “not found at compile time, at this URL” or it is not reported — is applied four times in this edition: to an independent evaluation of textGrain in item 01, to an independent evaluation of Beam in item 02, to the presidential action behind item 06, and to the lab desks above.

Previous
Previous

Edition 019 — 722 Math Manuscripts and the README's Fine Print

Next
Next

Edition 017 — Washington Builds Its AI Body, OpenAI Draws Its Line