AI this week
Last week the harness looked like the whole story: 37 points of spread on ARC-AGI-3, and we said so. This week a controlled study puts that number in its place. Across seven models and three harnesses on coding benchmarks, the harness barely moves whether the work succeeds — but it moves what it costs by a factor of two. The harness was never buying capability; it was buying, or wasting, money. That is the same story as last week, read one level down — and it is the level where a budget lives.
The three signals
01 — 03The harness is a cost line, not a capability line
Last week's 37-point spread came from two harnesses and two reasoning settings — we flagged the confound at the time. Arena.ai has now held everything else still, and the picture inverts: success barely moves, cost doubles. Which makes this a purchasing decision, currently being taken by default inside engineering teams.
2.0× Claude Code versus Pi, 1.6× versus Codex
97.8% / 96.7% / 96.7% — Claude Fable 5 across the three
$1.33 against $0.67 for the same work
The labs are writing the metrics they will be judged on
Three texts in seven days, from three labs, each proposing how frontier development should be measured. None of the figures is audited; the independent evaluators are announced, not seated. Whoever publishes a definition first tends to keep it — and it will reach you as a contract clause before it reaches you as a regulation.
~30,000 concurrent internal agents, 6% of AI R&D compute to safety
6 first misalignment disclosures at OpenAI, voluntary and self-declared
Employee-level access for third-party evaluators — pledged, not yet in place
European sovereignty consolidates by merger, not by raise
Mistral answered with €3B a fortnight ago. Cohere answers differently: it is buying the continuity of a legal framework rather than a place in the parameter race. For a European buyer under localisation requirements, this is the first credible option that does not route through an American hyperscaler.
Berlin + Toronto dual headquarters, Heidelberg kept as research
1,000+ staff, backed by Schwarz Group's STACKIT sovereign cloud
Closing expected later in 2026, subject to regulatory approvals
The rest of the week
20 entries| Player | Announcement | What to take from it |
|---|---|---|
| Arena.ai | HarnessTax: ±2% on success, up to 2.0× on cost | 21 model-harness pairs, 7 models, 3 harnesses (Claude Code, Codex CLI, Pi), 30 sampled tasks per benchmark run three times, at API prices dated 1 September 2026. A third-party measurement, not a vendor's. This is a cost line no executive committee tracks, and it corrects the read many drew from last week's ARC-AGI-3 spread. |
| Anthropic | Claude optimises 30+ biomolecular models, roughly 4× on average | Measured by Anthropic on NVIDIA H100 and H200, with near-2× at identical outputs on structure prediction and design; code and FlashPairformer kernels open-sourced. AI as performance engineer rather than researcher: the gain lands on the unit cost of scientific compute. If you run simulation at scale, the question is whether your own numerical code has been through this. |
| OpenAI | Astra for Law: 54.0% against 38.7% on Vals AI's Legal Research Bench validation set | GPT-6 Astra plus a legal index of 230M+ URLs covering 99.9% of published US precedential case law, via Trusted Access for selected firms; API announced as coming. The two scores compare a specialised model with an index against a general model with web search — not the same configuration. Vertical models are won on corpus access, not architecture. |
| Anthropic | Measurements of the pace of development inside the lab | Anthropic states Claude leads 26% of its AI R&D, up from under 1% in February 2026, with ~30,000 concurrent agents and about 6% of AI R&D compute going to safety. Nobody can verify this yet — which is why it is worth reading. The first published definition tends to become the standard the rest are held to. |
| Gemini 3.8 Live, 3.8 Live Extended Thinking and 3.5 Transcribe | $0.005 per minute of audio input and $0.018 on output, Google's own estimate based on $3 and $12 per million tokens; 85+ languages on Transcribe. Any contact-centre budget resting on a per-minute rate negotiated a year ago is now working from a stale assumption — reopen the price-review clause before reopening the architecture. |
| Player | Announcement | What to take from it |
|---|---|---|
| Anthropic | Cowork and chat merge into a single Claude | Claude Docs and Slides in beta, editable output, PowerPoint and PDF export; rolling out on Pro and Max over the coming weeks, Team and Free to follow. Enterprise admins get at least 30 days' notice — that window is when to check what the merge changes about retention and document sharing, not after. |
| Anthropic | Claude Code Projects redesigned around delegation | Claude scopes the request, delegates, coordinates parallel threads, reviews the outputs and assembles the result. Restricted beta on selected Pro and Max accounts, wider rollout announced. What is described is the shape of a team, not of a tool: decide now who signs off on the assembled result, and against what record. |
| OpenAI | Sponsored Agents in ChatGPT, with HubSpot and Shopify | A user who clicks an ad can open a clearly labelled conversation with a business-sponsored agent; selected US advertisers for now, Shopify app live for US merchants, international announced from 23 September. The point of conversion leaves your website for a conversation you only partly control — redefine the funnel metrics before committing budget. |
| Agent Substrate on GKE | Google announces 10× the density of standard container runtimes, sub-500ms resume, 500+ suspend/resume activations per second and 1,000+ dormant agents per host; open source for non-production, production support by allowlist. Vendor figures, unverified. When a dormant agent costs almost nothing to keep alive, the governing question becomes how many run without review. | |
| CC agent extended to groups of up to six | Shared calendars, tasks and admin forms across one household account, US only, 18+, waitlist; runs on Antigravity with Gemini models. Read it as a working prototype of multi-user governance over a shared inbox, calendar and documents — which is precisely the problem a project team or a service desk has. | |
| Meta | Meta One, from $2.99 to $499 a month | WhatsApp Plus at $2.99, Instagram and Facebook Plus at $3.99, individual bundles at $7.99 and $19.99, creator and business tiers from $14.99 to $499. Meta states 15 million subscriptions and trials to date, a figure that mixes the two. At $499 a month Meta is no longer selling a social add-on, it is pricing a professional tool. |
| Player | Announcement | What to take from it |
|---|---|---|
| OpenAI | Pre-IPO round reportedly discussed above $1.2tn | Bloomberg reports early talks, article paywalled, nothing signed and no confirmation from OpenAI; Altman called a 2026 listing "ill-advised" on 12 September. The market calendar slips while the valuation climbs — which means the funding stays private, and so does the accountability. Last week we flagged Anthropic's prospectus as the first verifiable set of numbers; this is the opposite move. |
| SoftBank | $11.9bn borrowed from about 20 banks | Above the $10bn first sought, with Son reportedly targeting close to $65bn into OpenAI by October; the share fell as much as 13% on Monday. The financing chain of the frontier now runs through bank debt secured against a single unlisted asset — map that link the way you would map a sole-source industrial supplier. |
| OpenAI | Glass Imaging reportedly acquired for more than $300m | Reported by the WSJ and relayed by TechCrunch, not confirmed by OpenAI; the company was founded in 2019 by two former Apple camera engineers. Buying computational optics is not diversification, it is preparing a sensor — vertical integration among AI vendors is reaching back into hardware, and away from anyone else's operating system. |
| NVIDIA | CUDA-Q extended to fault-tolerant quantum computing | CUDA-Q Logical, a shared design framework for programming logical rather than physical qubits, with partners including Infleqtion. Development tools, not an available machine. NVIDIA is locking the software layer before the hardware exists at scale, exactly as CUDA did for the GPU: long-range watching, not a budget decision this year. |
| Player | Announcement | What to take from it |
|---|---|---|
| Anthropic | Dario Amodei, "We Must Pace the Frontier" | A three-step plan: Anthropic unilaterally pledges third-party evaluators permanent employee-level access with the right to publish, then coordination between labs on capability-linked certification, then international agreements. No numerical thresholds. Zuckerberg and Altman both answered publicly within days. Certification by capability, embedded evaluator, right to publish — that is the vocabulary of your next procurement round. |
| OpenAI | Voluntary misalignment reporting framework, six first disclosures | Covers unauthorised model actions, safeguard failures and behaviour contradicting a published safety assessment; escalation to the Safety Advisory Group. OpenAI calls it an initial set, not a comprehensive account, and says it does not replace legal disclosure duties. Worth demanding in supplier due diligence — while noting it is voluntary, self-reported and explicitly incomplete. |
| Cohere / Aleph Alpha | Definitive business combination agreement | Operating as Cohere, dual-headquartered Berlin and Toronto, Heidelberg kept as a research centre, 1,000+ staff, with a STACKIT sovereign cloud partnership; Gomez stays CEO. Closing expected later in 2026, subject to final regulatory approvals, value undisclosed. Sovereignty consolidating by merger rather than by raise — but the deal has not closed. |
| Mistral | Open, private, multilingual AI in Firefox's Smart Window | A partnership with Mozilla, no deployment timetable or model scope disclosed. The browser is a contested distribution point again, and Mistral enters through privacy rather than performance — which gives a European buyer a defensible line in committee that has nothing to do with benchmark position. |
| Google DeepMind | Launch of the DeepMind Institute | Led by Hassabis, Manyika and Legg, studying the technical and societal implications of AGI across safety, governance, institutions and human values. All three major labs published a governance framework or body inside the same week — the regulatory debate is being drafted at the vendors before it reaches the regulators. |
Last week the harness looked free. This week it turns out you were paying for it twice
Seven days ago the harness arrived as a gift, and the honest question was what remained of the eighteen months you had spent building one. The answer is now sharper and less comfortable: the harness never bought capability, it bought cost — and the cheap one performs like the expensive one. Real-time voice has fallen below a cent a minute on input; a dormant agent costs almost nothing to keep alive. Each of those retires a budget assumption set less than a year ago. Meanwhile the vendors have begun writing the vocabulary of their own supervision, and that vocabulary will reach you as a contract clause long before it reaches you as a law.
- On our coding-agent work, which harness are we running, who chose it, and what does that choice cost against the cheapest alternative at comparable success rate?
- Do our AI supplier contracts carry a price-review clause triggered by a list-price cut, or are we paying a rate negotiated into a market that has since moved twice?
- What do we require from suppliers by way of safety measurement — and do we accept self-declared figures, or insist on a third-party evaluator with the right to publish, as Anthropic has pledged to host?
