AI

OpenAI declared the “AGI era.” The same model scores 62.7% on a neutral harness.

GPT-6 Astra posted 99.9% on ARC-AGI-3 using OpenAI’s own test rig. Run it through the provider-neutral one and it drops 37 points. Both numbers are real. Only one of them is yours.

N Noah · The Sharp Brief · September 4, 2026 · 4 min read

OpenAI shipped GPT-6 Astra on Wednesday and did something it has spent years carefully avoiding: it put a label on it. The launch materials frame the model as the beginning of the “AGI era,” and president Greg Brockman said he personally thinks that is what this is. For a company whose charter turns on the definition of that acronym, that is not marketing garnish. It is a position.

The number underneath the claim is ARC-AGI-3, the abstract-reasoning benchmark built specifically to be hard for models that have memorised the internet. Astra scored 99.9% on it. Its predecessor, GPT-5.6 Sol, scored 7.8%. A jump of ninety-two points in one generation on a benchmark designed to resist exactly that is, on its face, the most striking capability result of the year.

Then there is the harness. The 99.9% run used OpenAI’s own provider adapter, which retains reasoning state between turns and uses compaction to manage long contexts — so the score measures Astra plus OpenAI’s agent scaffolding, working as a system. Run the same model through the standard provider-neutral harness and it scores 62.7%. ARC Prize published both figures, and added the sentence that matters most: saturating the benchmark would not constitute proof of AGI. They call it meaningful progress toward generalisation. They do not call it AGI.

Our take: Both numbers are honest, and the gap between them is the actual product. 62.7% is what the weights do. 99.9% is what the weights do when someone has engineered the memory, the context compaction and the turn-to-turn state around them. OpenAI is no longer selling a model — it is selling a model welded to a harness, and it is benchmarking the weld. That is a defensible way to build a product and a terrible way to compare vendors, because your own scaffolding is not OpenAI’s. Assume you are buying the 62.7% and treat anything above it as work you have to do yourself.

What it costs, and what it can touch

Astra runs $10 per million input tokens and $50 per million output at the standard tier, with cached input at $1.00 and cache writes at $12.50 — a spread that makes cache discipline a line item rather than an optimisation. The context window is 1,050,000 tokens with 128,000 maximum output.

The more consequential detail is the tool surface. Astra ships with web search, a code interpreter, a hosted shell, apply-patch, computer use and MCP support. OpenAI is pitching it at filling forms, updating CRM records, doing online research, and building and QA-ing websites — a model that changes state in your systems rather than one that returns text about them. Access is rolling out first through the Trusted Access Program for enterprises, with API and Plus, Pro, Business and Enterprise plans following in the coming days. The staged rollout is itself the tell: the more a model acts, the less comfortable a lab is shipping it to everyone on day one.

What to watch

Advertisement

Get the day, decoded — at 7 PM ET

The Sharp Brief: AI, money, business & performance in five sharp minutes. Free.

Free bonus: subscribe today and The 2026 AI Playbook (PDF) lands with your welcome email.

Recommended by 5+ newsletters across AI, markets & business.