OpenAI shipped GPT-6 Astra on Wednesday and did something it has spent years carefully avoiding: it put a label on it. The launch materials frame the model as the beginning of the “AGI era,” and president Greg Brockman said he personally thinks that is what this is. For a company whose charter turns on the definition of that acronym, that is not marketing garnish. It is a position.
The number underneath the claim is ARC-AGI-3, the abstract-reasoning benchmark built specifically to be hard for models that have memorised the internet. Astra scored 99.9% on it. Its predecessor, GPT-5.6 Sol, scored 7.8%. A jump of ninety-two points in one generation on a benchmark designed to resist exactly that is, on its face, the most striking capability result of the year.
Then there is the harness. The 99.9% run used OpenAI’s own provider adapter, which retains reasoning state between turns and uses compaction to manage long contexts — so the score measures Astra plus OpenAI’s agent scaffolding, working as a system. Run the same model through the standard provider-neutral harness and it scores 62.7%. ARC Prize published both figures, and added the sentence that matters most: saturating the benchmark would not constitute proof of AGI. They call it meaningful progress toward generalisation. They do not call it AGI.
Our take: Both numbers are honest, and the gap between them is the actual product. 62.7% is what the weights do. 99.9% is what the weights do when someone has engineered the memory, the context compaction and the turn-to-turn state around them. OpenAI is no longer selling a model — it is selling a model welded to a harness, and it is benchmarking the weld. That is a defensible way to build a product and a terrible way to compare vendors, because your own scaffolding is not OpenAI’s. Assume you are buying the 62.7% and treat anything above it as work you have to do yourself.
What it costs, and what it can touch
Astra runs $10 per million input tokens and $50 per million output at the standard tier, with cached input at $1.00 and cache writes at $12.50 — a spread that makes cache discipline a line item rather than an optimisation. The context window is 1,050,000 tokens with 128,000 maximum output.
The more consequential detail is the tool surface. Astra ships with web search, a code interpreter, a hosted shell, apply-patch, computer use and MCP support. OpenAI is pitching it at filling forms, updating CRM records, doing online research, and building and QA-ing websites — a model that changes state in your systems rather than one that returns text about them. Access is rolling out first through the Trusted Access Program for enterprises, with API and Plus, Pro, Business and Enterprise plans following in the coming days. The staged rollout is itself the tell: the more a model acts, the less comfortable a lab is shipping it to everyone on day one.
What to watch
- Whether the neutral score moves. If independent harnesses converge upward toward the OpenAI number over the next few weeks, the scaffolding advantage is reproducible. If they stay near 62.7%, it is proprietary.
- Competitor harness disclosure. The pressure is now on every lab to say which rig produced which score. Whoever publishes both numbers first sets the norm.
- Cache economics in real workloads. A 10:1 gap between input and cached input means long-running agents are priced entirely by how well they reuse context.
- The charter question. OpenAI’s governance documents attach consequences to declaring AGI. Saying “AGI era” is not the same as saying “AGI,” and the distance between those phrases is legal, not linguistic.
