← All episodes
Special · AI News of the Week ·5:28 ·August 15, 2026

AI News of the Week — August 15, 2026

On August 7, OpenAI stopped part of the work on its next model — it could not rule out that the model writes and runs zero-days on hardened systems unsupervised. Three days later it shipped a model built to do most of that. Both calls were right. That's the week.

The Promise

  • Enterprise inference hit the lowest average of the year — Jefferies puts it at $1.16–$1.18 per million tokens, down from $2.04 in May, a fall of roughly 43%.
  • Meta's Muse Glimmer, around 30 billion parameters, quantises down onto a single 24GB consumer card. For a regulated environment where the blocker was always the prompt leaving the building, that's the first version of this that clears legal.
  • OpenAI stopped work on Astra when its own evaluations came back too strong — every activity the new controls did not yet cover halted, with benchmarking continuing. The preparedness framework did what it was written to do.
PROMISE RISK
Balanced

The Risk

  • The unit price fell while both OpenAI and Anthropic moved enterprise customers off flat subscriptions onto metered billing. Cheaper per token is not cheaper.
  • GPT-5.6 Cyber answers 95% of requests on OpenAI's internal advanced-security benchmark against the standard model's 1.5%. That benchmark is internal and nobody outside has checked it. Refusal stopped being the control.
  • Two days from 'not on the public API' to running on Amazon Bedrock. The gap between arrival and availability is the only defensive advantage anyone gets, and commercial pressure closes it every time.

Both calls were right

On 7 August OpenAI published a post about Astra, an unreleased model. Its own evaluations came back strong enough that, in its words, it cannot rule out critical capability level at this time — critical being the top of its preparedness framework. Every activity the new controls did not yet cover stopped. Benchmarking continued.

On 10 August it shipped GPT-5.6 Cyber through a gated programme called Daybreak Red: identity checks, signed legal attestations, and hardware security keys going mandatory for individual accounts on 1 September. Not on the public API.

On OpenAI’s own benchmark for advanced security work, this model answers 95% of requests. The standard model answers 1.5%. The last-generation cyber build answered 57. That benchmark is internal and nobody outside has checked it. What has been checked is a Chrome vulnerability the model found — reported in July, patched in July.

Then on 11 August AWS announced Daybreak Red and Daybreak Blue live on Amazon Bedrock for eligible customers; Bedrock’s model card dates the launch to 12 August. Two days from not on the public API to running on the largest cloud platform in the world.

I’ve watched this exact sequence for 25 years. A capability arrives, the control is that almost nobody can reach it, and then commercial pressure widens the door. Every time. The question was never whether attackers reach this class of tool — it’s what you did with the gap.

One regulator’s rule became everybody’s default

On 11 August Anthropic confirmed it now watermarks the text its models write. Models launched on or after 2 August carry it automatically; earlier ones sit in a transition period. Images and files get signed provenance under the open C2PA standard. In Anthropic’s own wording, it will travel with the text when copied and pasted elsewhere, and may persist through some editing.

The reason is Article 50 of the EU AI Act, live since 2 August, requiring machine-readable marking of generated content. But Article 50 reaches the European market — and Anthropic is marking worldwide. A response written for someone in São Paulo carries it, and no law asked for that.

Five days, five models — and almost no new pre-training

12 August: Grok 4.6, same base as 4.5, gains from post-training. 13 August was the crowded day — Gemini 3.7 Flash at half the price of the model it replaces; GPT-5.6 Small, ultra-fast at up to 750 tokens a second on Cerebras hardware; and DeepSeek taking V4 Pro out of preview with weights on Hugging Face, then telling customers prices rise as much as twelve-fold from the 17th. 14 August: Zhipu’s GLM 5.3 through its coding plan, weights about two weeks out after a security review.

Notice what almost none of these are: new pre-training runs.

Nothing this week changed what a model can do. Every move was about who gets to hold it. OpenAI answered with a partner list and a hardware key. Anthropic answered by marking what its new models write in countries with no law requiring it. Meta answered with a download link.

What’s left is access, provenance, and price — and all three are procurement questions. The ones who wait will buy the answer twice.