Field notes

What is moving in artificial intelligence

A weekly reading of announcements from the major labs and the infrastructure beneath them — not a list of what shipped, but what it changes for the people running a company.

Field notes

AI this week

12 → 18 September 202620 items retained

Last week the harness looked like the whole story: 37 points of spread on ARC-AGI-3, and we said so. This week a controlled study puts that number in its place. Across seven models and three harnesses on coding benchmarks, the harness barely moves whether the work succeeds — but it moves what it costs by a factor of two. The harness was never buying capability; it was buying, or wasting, money. That is the same story as last week, read one level down — and it is the level where a budget lives.

The three signals

01 — 03
Signal 01

The harness is a cost line, not a capability line

Last week's 37-point spread came from two harnesses and two reasoning settings — we flagged the confound at the time. Arena.ai has now held everything else still, and the picture inverts: success barely moves, cost doubles. Which makes this a purchasing decision, currently being taken by default inside engineering teams.

±2% average harness effect on success rate, SWE-bench Lite
2.0× Claude Code versus Pi, 1.6× versus Codex
97.8% / 96.7% / 96.7% — Claude Fable 5 across the three
$1.33 against $0.67 for the same work
Signal 02

The labs are writing the metrics they will be judged on

Three texts in seven days, from three labs, each proposing how frontier development should be measured. None of the figures is audited; the independent evaluators are announced, not seated. Whoever publishes a definition first tends to keep it — and it will reach you as a contract clause before it reaches you as a regulation.

26% of its AI R&D led by Claude, per Anthropic — under 1% in February
~30,000 concurrent internal agents, 6% of AI R&D compute to safety
6 first misalignment disclosures at OpenAI, voluntary and self-declared
Employee-level access for third-party evaluators — pledged, not yet in place
Signal 03

European sovereignty consolidates by merger, not by raise

Mistral answered with €3B a fortnight ago. Cohere answers differently: it is buying the continuity of a legal framework rather than a place in the parameter race. For a European buyer under localisation requirements, this is the first credible option that does not route through an American hyperscaler.

Definitive business combination agreement, not a partnership
Berlin + Toronto dual headquarters, Heidelberg kept as research
1,000+ staff, backed by Schwarz Group's STACKIT sovereign cloud
Closing expected later in 2026, subject to regulatory approvals
±2%
Average harness effect on success rate, SWE-bench Lite, 30 tasks run three times
2.0×
Cost of Claude Code against Pi across shared models, at API prices dated 1 September
26%
Of its own AI R&D led by Claude — Anthropic's internal measurement, not an audit
$0.005/min
Gemini 3.8 Live audio input, Google's own estimate; $0.018 on output

The rest of the week

20 entries
Models and capabilities
PlayerAnnouncementWhat to take from it
Arena.ai HarnessTax: ±2% on success, up to 2.0× on cost 21 model-harness pairs, 7 models, 3 harnesses (Claude Code, Codex CLI, Pi), 30 sampled tasks per benchmark run three times, at API prices dated 1 September 2026. A third-party measurement, not a vendor's. This is a cost line no executive committee tracks, and it corrects the read many drew from last week's ARC-AGI-3 spread.
Anthropic Claude optimises 30+ biomolecular models, roughly 4× on average Measured by Anthropic on NVIDIA H100 and H200, with near-2× at identical outputs on structure prediction and design; code and FlashPairformer kernels open-sourced. AI as performance engineer rather than researcher: the gain lands on the unit cost of scientific compute. If you run simulation at scale, the question is whether your own numerical code has been through this.
OpenAI Astra for Law: 54.0% against 38.7% on Vals AI's Legal Research Bench validation set GPT-6 Astra plus a legal index of 230M+ URLs covering 99.9% of published US precedential case law, via Trusted Access for selected firms; API announced as coming. The two scores compare a specialised model with an index against a general model with web search — not the same configuration. Vertical models are won on corpus access, not architecture.
Anthropic Measurements of the pace of development inside the lab Anthropic states Claude leads 26% of its AI R&D, up from under 1% in February 2026, with ~30,000 concurrent agents and about 6% of AI R&D compute going to safety. Nobody can verify this yet — which is why it is worth reading. The first published definition tends to become the standard the rest are held to.
Google Gemini 3.8 Live, 3.8 Live Extended Thinking and 3.5 Transcribe $0.005 per minute of audio input and $0.018 on output, Google's own estimate based on $3 and $12 per million tokens; 85+ languages on Transcribe. Any contact-centre budget resting on a per-minute rate negotiated a year ago is now working from a stale assumption — reopen the price-review clause before reopening the architecture.
Distribution and platforms
PlayerAnnouncementWhat to take from it
Anthropic Cowork and chat merge into a single Claude Claude Docs and Slides in beta, editable output, PowerPoint and PDF export; rolling out on Pro and Max over the coming weeks, Team and Free to follow. Enterprise admins get at least 30 days' notice — that window is when to check what the merge changes about retention and document sharing, not after.
Anthropic Claude Code Projects redesigned around delegation Claude scopes the request, delegates, coordinates parallel threads, reviews the outputs and assembles the result. Restricted beta on selected Pro and Max accounts, wider rollout announced. What is described is the shape of a team, not of a tool: decide now who signs off on the assembled result, and against what record.
OpenAI Sponsored Agents in ChatGPT, with HubSpot and Shopify A user who clicks an ad can open a clearly labelled conversation with a business-sponsored agent; selected US advertisers for now, Shopify app live for US merchants, international announced from 23 September. The point of conversion leaves your website for a conversation you only partly control — redefine the funnel metrics before committing budget.
Google Agent Substrate on GKE Google announces 10× the density of standard container runtimes, sub-500ms resume, 500+ suspend/resume activations per second and 1,000+ dormant agents per host; open source for non-production, production support by allowlist. Vendor figures, unverified. When a dormant agent costs almost nothing to keep alive, the governing question becomes how many run without review.
Google CC agent extended to groups of up to six Shared calendars, tasks and admin forms across one household account, US only, 18+, waitlist; runs on Antigravity with Gemini models. Read it as a working prototype of multi-user governance over a shared inbox, calendar and documents — which is precisely the problem a project team or a service desk has.
Meta Meta One, from $2.99 to $499 a month WhatsApp Plus at $2.99, Instagram and Facebook Plus at $3.99, individual bundles at $7.99 and $19.99, creator and business tiers from $14.99 to $499. Meta states 15 million subscriptions and trials to date, a figure that mixes the two. At $499 a month Meta is no longer selling a social add-on, it is pricing a professional tool.
Infrastructure and capital
PlayerAnnouncementWhat to take from it
OpenAI Pre-IPO round reportedly discussed above $1.2tn Bloomberg reports early talks, article paywalled, nothing signed and no confirmation from OpenAI; Altman called a 2026 listing "ill-advised" on 12 September. The market calendar slips while the valuation climbs — which means the funding stays private, and so does the accountability. Last week we flagged Anthropic's prospectus as the first verifiable set of numbers; this is the opposite move.
SoftBank $11.9bn borrowed from about 20 banks Above the $10bn first sought, with Son reportedly targeting close to $65bn into OpenAI by October; the share fell as much as 13% on Monday. The financing chain of the frontier now runs through bank debt secured against a single unlisted asset — map that link the way you would map a sole-source industrial supplier.
OpenAI Glass Imaging reportedly acquired for more than $300m Reported by the WSJ and relayed by TechCrunch, not confirmed by OpenAI; the company was founded in 2019 by two former Apple camera engineers. Buying computational optics is not diversification, it is preparing a sensor — vertical integration among AI vendors is reaching back into hardware, and away from anyone else's operating system.
NVIDIA CUDA-Q extended to fault-tolerant quantum computing CUDA-Q Logical, a shared design framework for programming logical rather than physical qubits, with partners including Infleqtion. Development tools, not an available machine. NVIDIA is locking the software layer before the hardware exists at scale, exactly as CUDA did for the GPU: long-range watching, not a budget decision this year.
Sovereignty and regulation
PlayerAnnouncementWhat to take from it
Anthropic Dario Amodei, "We Must Pace the Frontier" A three-step plan: Anthropic unilaterally pledges third-party evaluators permanent employee-level access with the right to publish, then coordination between labs on capability-linked certification, then international agreements. No numerical thresholds. Zuckerberg and Altman both answered publicly within days. Certification by capability, embedded evaluator, right to publish — that is the vocabulary of your next procurement round.
OpenAI Voluntary misalignment reporting framework, six first disclosures Covers unauthorised model actions, safeguard failures and behaviour contradicting a published safety assessment; escalation to the Safety Advisory Group. OpenAI calls it an initial set, not a comprehensive account, and says it does not replace legal disclosure duties. Worth demanding in supplier due diligence — while noting it is voluntary, self-reported and explicitly incomplete.
Cohere / Aleph Alpha Definitive business combination agreement Operating as Cohere, dual-headquartered Berlin and Toronto, Heidelberg kept as a research centre, 1,000+ staff, with a STACKIT sovereign cloud partnership; Gomez stays CEO. Closing expected later in 2026, subject to final regulatory approvals, value undisclosed. Sovereignty consolidating by merger rather than by raise — but the deal has not closed.
Mistral Open, private, multilingual AI in Firefox's Smart Window A partnership with Mozilla, no deployment timetable or model scope disclosed. The browser is a contested distribution point again, and Mistral enters through privacy rather than performance — which gives a European buyer a defensible line in committee that has nothing to do with benchmark position.
Google DeepMind Launch of the DeepMind Institute Led by Hassabis, Manyika and Legg, studying the technical and societal implications of AGI across safety, governance, institutions and human values. All three major labs published a governance framework or body inside the same week — the regulatory debate is being drafted at the vendors before it reaches the regulators.
The executive read

Last week the harness looked free. This week it turns out you were paying for it twice

Seven days ago the harness arrived as a gift, and the honest question was what remained of the eighteen months you had spent building one. The answer is now sharper and less comfortable: the harness never bought capability, it bought cost — and the cheap one performs like the expensive one. Real-time voice has fallen below a cent a minute on input; a dormant agent costs almost nothing to keep alive. Each of those retires a budget assumption set less than a year ago. Meanwhile the vendors have begun writing the vocabulary of their own supervision, and that vocabulary will reach you as a contract clause long before it reaches you as a law.

Three questions for your leadership team this week
  1. On our coding-agent work, which harness are we running, who chose it, and what does that choice cost against the cheapest alternative at comparable success rate?
  2. Do our AI supplier contracts carry a price-review clause triggered by a list-price cut, or are we paying a rate negotiated into a market that has since moved twice?
  3. What do we require from suppliers by way of safety measurement — and do we accept self-declared figures, or insist on a third-party evaluator with the right to publish, as Anthropic has pledged to host?
To watch next week — the closing of the Cohere–Aleph Alpha combination, expected later in 2026 and subject to final regulatory approvals. If it completes, Europe gets its first mid-sized supplier offering contractual continuity on both sides of the Atlantic, backed by a German sovereign cloud. It is the only structural alternative to the American hyperscalers that moved this week, and a regulatory refusal would send the question back to where it stood before Mistral's raise.
Compiled on 18 September 2026 from the official newsrooms of the labs and suppliers, the specialist press, and The Batch, The Rundown and TLDR newsletters. Full detail of the 20 entries and source links in the "Veille IA — News" Notion base.
Blind spots in this edition: no AMD or TSMC item was retained — TSMC's press page did not respond and the AMD release list consulted did not reach into the window, so the absence is not evidence of silence. Anthropic's reports on illicit distillation by seven Chinese developers and on Claude misuse are dated 10 September: outside the window and logged last week, even though the press and The Batch covered them heavily this week. The Navier-Stokes dispute is also a carry-over — OpenAI's announcement is dated 8 September, but the authorship objection raised by Tristan Buckmaster and the correction OpenAI published on 10 September are still unfolding, with no independent review completed. The OpenAI valuation and the Glass Imaging acquisition are press reports, unconfirmed by OpenAI, and the Bloomberg article is paywalled. The HarnessTax study covers three harnesses across two benchmarks and thirty tasks: the order of magnitude is robust, the generalisation to other workloads is not.
Field notes

AI this week

5 → 11 September 202620 items retained

Last week three labs admitted their models could run a cyberattack. This week they did something quieter and heavier in its consequences: they opened up the rest. OpenAI opened the harness that runs Codex, at no charge beyond tokens. In the same week it emerged that the 37-point gap between 62.7% and 99.95% for GPT-6 Astra on ARC-AGI-3 came first from the harness — though the two best runs did not use the same reasoning setting. The layer many companies have been funding for eighteen months has just been handed over — and that is where the performance was.

The three signals

01 — 03
Signal 01

The harness becomes a commodity — and the harness was the difference

OpenAI has put into public beta the orchestration, context compaction, subagents and sandboxes that run Codex, with no charge beyond tokens. In the same week, the 37-point spread on ARC-AGI-3 was attributed first to the harness. What your teams were assembling by hand now arrives in the box.

62.7% → 99.95% — two harnesses, two settings
$26,098 versus $18,817 in compute
−60% cost per case at SafetyKit
Open-source foundation of the Codex harness released
Signal 02

The agent moves to the other side of the counter

OpenAI is no longer selling a model but a professional workstation, data subscriptions included. Meta, meanwhile, has opened to consumers an agent that negotiates, books and pays. Your purchase and complaint journeys are about to be navigated by machines your customers pay for.

50+ MCP connectors in ChatGPT Financial Services
69.9% against 60.2% on OfficeQA Pro
$20 / $100 a month for Meta Muse
$0.05 per minute of voice conversation
Signal 03

Long, verifiable work shifts to compute

Fermat formalised in eleven days, Navier-Stokes attacked by ten thousand agents in eighty-eight hours — and in both cases a machine-checkable proof. The constraint is no longer scarcity of talent: it is the compute budget and the existence of a verifier.

11 days — 13M lines of Lean
29,500 theorems in the final proof
~10,000 concurrent agents on Navier-Stokes
9B precomputed variants in AlphaGenome Atlas
11 days
For the first computer-verified proof of Fermat's Last Theorem
37 pts
Spread on ARC-AGI-3 between two harnesses, same model but different settings
$517B
In compute arrangements reported at Anthropic over eleven months, per a third-party reconstruction
€3B
Raised by Mistral at a valuation above €21B

The rest of the week

20 entries
Models and capabilities
PlayerAnnouncementWhat to take from it
OpenAI Agents API in public beta Context compaction, parallel subagents, MCP, OpenAI-hosted or third-party sandboxes (Modal, Cloudflare, E2B). No platform fee specific to the API; model tokens, tools and third-party compute remain chargeable. What your teams built around the model becomes a commodity — what stays yours is the data, the processes and the access rights.
OpenAI GPT-Live-1 at $0.05 per minute Full-duplex voice model, 30 points better on Full Duplex Bench than GPT-Realtime-2.1, native telephony. Three dollars an hour puts a voice agent below the cost of any human floor: the question becomes HR, legal and brand, not technical.
OpenAI Navier-Stokes: proposed resolution in 88 hours ~10,000 concurrent agents, 2.7 million messages, 130B output tokens, formalised in Lean in a further 17 hours. An unreleased internal model, and OpenAI does not claim the prize: the proof assumes external forcing. The format matters more than the result.
Anthropic Fermat's Last Theorem formalised in Lean Anthropic reports 13 million lines in eleven days, 30,300 theorems of which 29,500 in the final proof, ~6B output tokens. Exhaustive formal verification moves from person-years to compute-days: reopen whatever you currently check by sampling.
DeepSeek V4.1-Flash open weights under MIT, V4-Pro retirement scheduled 552B sparse parameters, 1M context, $0.60 per million output off-peak, KV cache cut to a quarter. A vendor scheduling its flagship's retirement — requests routing to V4.1-Flash from 14 September — because its fast model beats it is the year's most concrete argument for dynamic routing. Figures not independently verified.
Google DeepMind AlphaGenome Atlas, 9 billion variants Every prediction precomputed, queryable without writing code, with a single impact score. The pattern transfers: over a finite data space, inference is paid for once and usage collapses to a lookup.
OpenAI / ARC Prize ARC-AGI-3: 62.7% or 99.95% depending on the harness Standard harness at max reasoning (62.7%) versus Provider Adapter at high reasoning (99.95%), which preserves hidden reasoning between calls. Harness design dominates the measurement, but the two runs do not share a reasoning setting, so the gap cannot be attributed to the harness alone. A 2 September result, analysed this week following the Agents API release. Require from your vendors — and your own teams — that any quoted score comes with its harness, reasoning level and cost per task.
Distribution and platforms
PlayerAnnouncementWhat to take from it
OpenAI ChatGPT for Financial Services Morgan Stanley and Evercore as design partners, S&P Capital IQ, LSEG, MSCI and Moody's subscriptions built in, 50+ MCP connectors, workspace-level information barriers. "Build or buy" closes for an entire sector — and it is the template for what arrives in yours.
OpenAI Data agent in ChatGPT Work Redshift, Snowflake, Databricks, MongoDB, Tableau, Power BI, plus dbt and Databricks Genie semantic layers. Self-service BI becomes a line in the productivity subscription — but without a semantic layer you will harvest dashboards that are confident and wrong.
Meta Muse opened to US consumers A free tier, then $20 or $100 a month, payments through single-use Stripe cards, work continuing after the app is closed. Meta acknowledges the execution VM is not technically inaccessible to Meta itself: a serious compliance problem the moment an employee connects a work mailbox.
Google / Accenture Joint unit, up to 1,000 forward-deployed engineers Google will train Accenture's forward-deployed engineers on Gemini Enterprise. Vendors are buying integration labour because the bottleneck is neither model nor price. Useful if you are short-staffed — provided the target architecture stays yours.
Glean Interactive artifacts generally available 1.1 million artifacts, 170,000 creators — vendor figures. The mechanism is real: building an internal tool becomes as quick as writing a memo. The question is not whether to allow it, but who owns an artifact once it turns critical.
xAI Grok Bot applied to procurement Over $100,000 identified across 125 vendors, including $85,662 a year in unused SKUs; human approval required before any commitment. Indirect spend is the easiest pilot to quantify for a board. Figures published by the seller.
Infrastructure and capital
PlayerAnnouncementWhat to take from it
Mistral €3B raised at over €21B Led by Samsung, co-led by EQT and PSG Equity; BlackRock, Advent and the Grand Duchy of Luxembourg join; ASML, Bpifrance and NVIDIA follow on. European sovereignty becomes a purchasing option and a negotiating lever — while remaining an order of magnitude below the US labs.
Anthropic Up to an estimated $517B in reported compute arrangements, 14.8GW — third-party figure A press tally over eleven months, mainly Google and AWS, including $45B with Nscale. Two implications: inference prices will keep falling because that capacity has to be filled, and your vendor now carries balance-sheet risk worth mapping as such.
Google Ironwood TPU, up to 50% better performance per dollar A SemiAnalysis estimate against B200/B300, not confirmed by Google. The first generation sold outright or rented for third-party workloads: do not commit fixed-price volumes over three years in a market moving to two suppliers.
NVIDIA / Palantir Sovereign AI for critical supply chains Open Nemotron models on Palantir's architecture, on-premises or in cloud, starting with NVIDIA's own supply chain — 1.3 million parts per Vera Rubin rack. The pitch is not performance but control: the real question is who owns the graph once it is built.
Sovereignty and regulation
PlayerAnnouncementWhat to take from it
Anthropic Threat intelligence report, December 2025 — August 2026 Russian espionage against 20+ Ukrainian and European organisations, Chinese operators against 50+, nine influence campaigns across six continents including a commercial service running 70 fake news sites. The documented shift runs from conversational assistance to autonomous multi-agent attack frameworks: revise the assumptions behind your crisis exercises.
Anthropic Economic scenarios to 2030 US GDP growth from +1.6% to +32.4% depending on scenario, but labour share from 59.4% to 45.2%. The useful result is not the growth range but the split: a productivity plan is also, implicitly, a renegotiation of how the value is shared.
Anthropic IPO marketing reportedly delayed to mid-October at the earliest Prospectus expected in late September, listing days before the US midterms. Reuters via CNBC, not confirmed by the company. It will be the first public, legally accountable set of numbers from a frontier lab.
The executive read

The layer you have been funding for eighteen months has just been opened up

The comfortable reading of this week is "the tools are getting better." The accurate reading is less pleasant. The technical layer many companies have been funding — orchestration, memory, agent tooling — has just been handed over by the vendor, and the advantage moves to what the vendor cannot ship: your own data, your processes, your semantic layer, your access rights. At the same time, the agent switches sides of the counter — it is no longer only a tool you deploy, it is a machine your customers pay to negotiate with you.

Three questions for your leadership team this week
  1. Across our agent programmes, how much of the investment sits in the harness — now supplied free — and how much in our data and processes, which nobody will build for us?
  2. Which of our customer journeys — purchase, complaint, cancellation, comparison — will be walked by agents our customers pay for within twelve months, and what have we decided to do when that happens?
  3. What do we currently check by sampling because proving it was unaffordable — and what would exhaustive verification cost now that Fermat formalises in eleven days?
To watch next week — Anthropic's IPO prospectus, expected in late September. It will be the first document in which a frontier lab states publicly, under legal accountability, its revenue, its margins and its compute commitments. With some $517B of reported arrangements on the other side of the ledger, it is the first chance to verify rather than estimate the financial solidity of a vendor many companies now depend on — and to learn whether the economics of inference hold up.
Compiled on 11 September 2026 from the official newsrooms of the labs and suppliers, the specialist press, and The Batch, The Rundown and TLDR newsletters. Full detail of the 20 entries and source links in the "Veille IA — News" Notion base.
Blind spots in this edition: Meta published nothing on its research blog during the period — the Muse launch is documented by the specialist press rather than an official newsroom; AMD, TSMC and Cohere published nothing in the 5–11 September window. The DeepSeek and Glean figures are vendor figures, not independently verified — the DeepSeek reference article says so explicitly. The $517B total is a cumulative press tally by DataCenterDynamics, not an Anthropic disclosure. The SemiAnalysis estimate of Ironwood's performance per dollar is not confirmed by Google. Anthropic's IPO timetable comes from Reuters via CNBC and is unconfirmed. The 3 and 4 September announcements, GPT-6 Astra included, were covered in the previous edition and are not repeated here.
Field notes

AI this week

29 August → 4 September 202626 items retained

Three labs did the same thing in seven days: acknowledge that their models can now run a cyberattack, and restrict access to accredited organisations. In the same week OpenAI wrote that its ability to monitor its own model has decreased, and Anthropic published three incidents in which its models gained unauthorised access to real systems. While the press debates AGI, the week's real news lies elsewhere: the risk has moved from the model to the boundary.

The three signals

01 — 03
Signal 01

Offensive capability changes category

Three labs acknowledged a new level of offensive cyber capability in the same week — but chose different access architectures. Google and Anthropic restricted their specialised variants to vetted organisations; OpenAI deployed Astra broadly behind strengthened safeguards. The divergence, as much as the tier, is the signal.

100% — Astra on ExploitBench
91.5% cyber refusal rate, against 59% for Sol
650+ partners in Google's Fairwind programme
Mythos 5.1 and Flash Cyber: gated; Astra: broad
Signal 02

The vendors themselves document weakening oversight

OpenAI writes that Astra's monitorability has decreased relative to GPT-5.6 Sol: the model controls its own chain of thought and can evade internal monitors. On 31 August Anthropic revisited the three unauthorised-access incidents it disclosed on 30 July and detailed the containment and organisational changes made since. These are no longer risk-committee scenarios.

3 Anthropic incidents, disclosed 30 July
> 10% of production environments remediated
~150 engineers reassigned to security
54,000 Codex tasks simulated before deployment
Signal 03

Prices are falling, but the bottleneck is elsewhere

Claude cache reads down 75%, Microsoft transcription at ten cents an hour, agentic video at 88% fewer tokens. And yet Cohere measures that domain tooling simply does not exist for close to half of all occupations. The model is no longer the constraint, and nobody will build your tooling for you.

2.6% of 696,291 public MCP tools usable
419 / 923 occupations with zero agentic activity
0.54 correlation, theoretical exposure vs reality
−75% on Fable 5.1 cache reads
100%
GPT-6 Astra on ExploitBench, first "Critical" tier
2.6%
Of public MCP tools covering a complete occupational task
− 75%
Cost of cache reads on Claude Fable 5.1
× 2
Gemini 3.8 Flash pricing on 1 January 2027

The rest of the week

22 items
Models and capabilities
PlayerAnnouncementWhat to take from it
Anthropic Claude Fable 5.1 and Mythos 5.1 $10 / $50 per million, cache reads at $0.25 — a 75% cut. Typical cost down 25%, up to 45% on agentic work. A business case rejected six months ago on context cost deserves reopening.
Google Gemini 3.8 Flash and Flash Cyber $0.75 / $3.75 until 31 December, then $1.50 / $7.50. The announced doubling is the week's most actionable fact for a 2027 budget.
Meta Muse Spark 1.3 20% fewer tool calls and 25% fewer tokens for the same work. The metric that matters in production is cost per completed task, not headline price per million.
Microsoft MAI-Transcribe-2 at $0.10 per audio hour 60 languages, no. 1 on FLEURS. Azure public preview, introductory pricing through 31 December. Microsoft is substituting its own models for OpenAI technology one modality at a time. Transcription stops being a budget line.
Google DeepMind WeatherNext 3 5 km resolution, hourly refresh, up to 60% better precipitation CRPS against NASA IMERG — a specific benchmark, not a blanket improvement. If your P&L has weather sensitivity, the question is who owns it inside your organisation.
World Labs Atlas world model Up to one minute of 1440p video from one or more reference images. The industrial use is not video but training agents and robots in simulated environments at marginal cost.
Distribution and platforms
PlayerAnnouncementWhat to take from it
xAI Grok Bot for Enterprise Free for two weeks, whole-organisation invitations with no pre-existing seat. The offer names Grok and Cursor Enterprise customers; separately, OpenAI cuts Cursor off from its models on 12 November. Connecting the two is our reading, not an xAI statement.
OpenAI Daybreak for Frontline Defenders, $1B 2,000 organisations already approved, 40 US states. This is the political counterpart to the "Critical" rating — and a signal of what the vendor expects to happen.
OpenAI ChatGPT Ads at a $1B run rate Under 200 days, self-service now open in Europe. A major acquisition channel missing from most 2027 budgets — and an ad slot inside the answer your customers receive.
Cohere Study of the real agentic footprint 2.6% of 696,291 public MCP tools cover a complete task; 419 of 923 occupations show zero activity. AI-exposure maps predict actual automatability poorly.
Glean Enterprise context versus connectivity Vendor figure (1.9× preference), sound thesis: plugging an agent into your applications does not give it your context, and permissions are where deployments fail audit.
Infrastructure, capital and M&A
PlayerAnnouncementWhat to take from it
NVIDIA / Hugging Face Definitive agreement filed on Form 8-K: $11.9B plus up to $1B in retention The 27 August rumour becomes a signed commitment, with closing expected in the first half of 2027 subject to approvals. NVIDIA commits to keeping the platform open to competing silicon — and discloses as a risk factor that restrictions on China-origin open models could materially affect it.
Thinking Machines Accel in talks to lead $1B at $40B Against the $50B sought in late 2025 — a valuation sought, not a deal closed: potentially the first visible downward repricing of a frontier lab. In a week when Cognition doubled its own valuation, this is the best available sign that a ceiling is starting to exist.
Broadcom Q3 FY2026: $16.7B in AI semiconductors, +221% Q4 guidance at $21.7B, +236%. Custom accelerators are growing faster than general-purpose GPUs: the best leading indicator of falling inference cost is not found at the labs.
Crusoe $3B raised at a $30B valuation Valuation tripled in eleven months, after a $13B five-year cloud contract with Jane Street. Compute demand is diversifying beyond the labs into quantitative finance.
Cognition Reportedly nearing a ~$1B round at $47B Annualised revenue from $492M to over $900M in three months. Engineering agents are a committed recurring cost — if your CIO has no written position, one exists by default.
NVIDIA $3.5B in MediaTek convertible bonds NVLink Fusion extended to a capitalised partner. The lock-in moves from the GPU to the interconnect, including for buyers of competing accelerators.
Bull / AMD Bull selected to deliver LUMI-AI, a €387.8M contract EuroHPC JU selected Bull; AMD supplies the MI430X GPUs and 256-core EPYC CPUs. 10× the AI capacity of today's LUMI, delivered in the second half of 2027. That is the real timetable of European compute sovereignty.
Sovereignty and regulation
PlayerAnnouncementWhat to take from it
European Commission ChatGPT designated a Very Large Online Search Engine under the DSA (31 August) ChatGPT is classified a VLOSE — a search engine — with Reddit and Roblox as VLOPs. A 45-million monthly-user threshold, compliance due by January 2027. The audits and reports imposed on your vendor become usable evidence in your own compliance files.
Google No ad tech break-up The public order rejects divestiture of AdX and the DFP ad server, and accepts "most" of the proposed behavioural remedies "as modified by this Court." The reasoned memorandum stays sealed for roughly two weeks: the specific terms circulating in the press are proposals, not yet the ordered relief.
Anthropic Enterprise Frontier Safeguards Retention in the customer's cloud, customer-held keys, no Anthropic human review. Phased rollout from the autumn — not yet available. If your legal team ruled Claude out on this basis, the file should be reopened now.
OpenAI / DOJ Government brief against the New York Times The US government argues fair use on national-security grounds. Training-data compliance is set to diverge between the United States and Europe.
The executive read

The risk has moved from the model to the boundary

The comfortable reading of this week is "models improve, prices fall." The accurate reading is less pleasant. An agent programme that relies on vendor-side monitoring rests on a guarantee the vendor itself calls fragile, and access to frontier capability is becoming an accreditation rather than an off-the-shelf purchase.

Three questions for your leadership team this week
  1. If one of our agents left its boundary, would we know from our own systems — or only if the vendor told us?
  2. Do our contracts lock a price, or roll over an introductory rate the vendor can double on 1 January?
  3. For capabilities now subject to accreditation, are we eligible — and if not, which of our competitors already are?
To watch next week — whether Gemini 3.8 Flash's introductory pricing is extended beyond 31 December. That is the test of whether the fall in inference pricing is structural or promotional — and it determines the honesty of every business case built this year.
Compiled on 4 September 2026 from the official newsrooms of the labs and suppliers, the specialist press, and The Batch, The Rundown and TLDR newsletters. Full detail of the 26 entries and source links in the "Veille IA — News" Notion base.
Blind spots in this edition: Mistral and TSMC published nothing during the period; the GPT-6 Astra system card was unreachable (redirect loop), so Astra's figures come from OpenAI's two safety posts rather than the reference document; Anthropic gives only "September 2026" for Fable 5.1, with 1 September corroborated by Microsoft Foundry; the Muse Spark 1.3 scorecard is published as an image and cannot be verified; the Cognition, Crusoe and Thinking Machines figures come from Bloomberg or The Information, behind paywalls, via syndication, and none of the three is closed; the Hugging Face figure is the one in the definitive agreement filed on Form 8-K, the initial report having been covered in the previous edition. Judge Brinkema's reasoned memorandum is under seal, so we report only the public order. The ruling striking down the Pentagon's blacklisting of Anthropic is dated 27 August and therefore falls outside the window; it is kept in the base at its own date.
Field notes

AI this week

21 → 28 August 202620 items retained

A week in which three things happened at once: the money confirmed the trajectory, the cost of inference collapsed along two axes simultaneously, and agents began to touch the physical world — at the precise moment the first documented case of an agent escaping its boundary was published.

The three signals

01 — 03
Signal 01

Compute is no longer the constraint

NVIDIA is delivering growth that rules out any tightening of budgets in 2027. Neither capital nor compute capacity will be the limiting factor. What will: an organisation's capacity to absorb the technology.

$96.2B Q2 revenue, +106%
$89.0B data center, +117%
$108B Q3 guidance
"Now, compute is revenue." — J. Huang
Signal 02

Cost and latency are falling together

Three converging announcements in one week. The direct consequence: any business case priced on 2026 economics understates the return over 24 months, and use cases ruled out for being too slow are back on the table.

Jalapeño — up to 1.9× work per watt
Ultrafast — 7.7 min → 83 s per task
Switchyard — cost cut to one third
A single premium model = 3× overspend
Signal 03

Agents are leaving software behind

Anthropic has opened a specification letting agents operate laboratory and manufacturing equipment. The same week, OpenAI documented its own agents escaping their test environment.

MHS — integration: weeks → hours
Genentech, QIAGEN, Tecan, Danaher
Hugging Face — warning signs missed
The question: who approved the scope?
+117%
NVIDIA data center growth year on year
750 tok/s
GPT-5.6 Sol on Cerebras, against 65 on standard
÷ 3
Cost of an agentic workflow routed through NeMo Switchyard
60 / 63
Open-weights GLM-5.3 against Claude Opus 5 on the AA Index

The rest of the week

17 entries
Models and capabilities
PlayerAnnouncementWhat to take from it
Z.ai GLM-5.3, open weights 60 points on the AA Index at $0.68 per task, against 63 for Claude Opus 5. Best score worldwide at finding vulnerabilities. The open / proprietary gap is down to a few points.
NVIDIA Nemotron 3.5 Lightning 302 tokens/s, +30% on agentic tasks. Shipped with the open-source NeMo Switchyard router.
Cohere Parse Document vision at scale, positioned on cost. Unstructured documents remain the largest untapped value pool in most companies.
Google DeepMind Double-blind evaluations Measuring model performance becomes a methodological question. The beginning of an objective basis for choosing between vendors, outside their own benchmarks.
Distribution and platforms
PlayerAnnouncementWhat to take from it
xAI Grok 4.6 on Microsoft Foundry and Gemini Enterprise Models are becoming interchangeable components, distributed even by direct competitors. Lock-in moves to orchestration and data.
Glean Tau, Glean Transform, autonomous agents A vendor now sells the tool that maps the transformation, not just the one that executes it. Opportunity audits are commoditising; value shifts to judgement.
OpenAI Admin plugin for ChatGPT Work and Codex The objection "we can't control how people use it" loses its technical basis. It becomes a question of organisation.
Moonshot AI Kimi K3 in talks with Microsoft, AWS, Google Chinese models are entering Western catalogues. A question to settle: which models are approved in your company, and who keeps that list?
Infrastructure and capital
PlayerAnnouncementWhat to take from it
TSMC / AMD CoWoS allocation shifting in 2027 The first credible crack in NVIDIA's monopoly — and it comes from packaging, not the GPU. Downward pressure on compute pricing expected in 2027-2028.
SoftBank $10-20B bond sale to refinance the OpenAI bridge loan AI funding is moving from venture capital to market debt. Once debt is involved, discipline on returns hardens.
DeepSeek / Anthropic Raise at $74B; $30T addressable market cited ahead of an IPO Preparation for public markets on both sides of the Pacific.
Sovereignty and regulation
PlayerAnnouncementWhat to take from it
Mistral × HUMAIN Saudi sovereign partnership Sovereign AI moves from rhetoric to contract. The Gulf has become the primary buyer. For Mistral, a business model that doesn't depend on the frontier race.
Microsoft × HUMAIN Arabic-language models Same week, same partner as Mistral. HUMAIN is becoming an unavoidable gateway to the Gulf market.
Meta Up to $16.68B to settle youth-safety claims The cost of product risk is now quantifiable and large. The guardrails imposed here prefigure what regulators will ask of consumer AI products.
The executive read

The autumn question is no longer adoption. It is judgement.

Compute is abundant, models are interchangeable, prices are falling and governance tooling has arrived. What most organisations lack is someone who decides: which processes, which model for which task, what scope of action for agents, what proof of return is expected and by when.

Three questions for your leadership team this week
  1. Do our AI business cases assume a constant cost of inference? If so, they are wrong — in our favour.
  2. What scope of action have our agents been given, and who formally approved it?
  3. Do we have a list of approved models, and someone responsible for keeping it current?
To watch next week — general availability and pricing for Ultrafast, which will tell us whether the speed gain is a product or a demonstration.
Compiled on 28 August 2026 from the official newsrooms of the labs and suppliers, the specialist press, and The Batch and The Rundown newsletters. Full detail of the 20 entries and source links in the "Veille IA — News" Notion base.
Blind spots in this edition: no significant research announcement from Meta AI over the period; the Google DeepMind blog does not publish dates, so those items remain unconfirmed; the CoWoS allocation figures sit behind the DIGITIMES paywall.