AI

Anthropic’s new model took the top spot by one point. Its hallucination rate moved 14.

Claude Opus 5 scores 61 on the Artificial Analysis Intelligence Index — one point above Anthropic’s own Fable 5, two above GPT‑5.6 Sol, four above Kimi K3. It is cheaper per task than Fable, though by 26%, not the half the token prices imply. And on the factual-recall test it now answers when it doesn’t know.

N Noah · The Sharp Brief · July 26, 2026 · 4 min read
Abstract benchmark leaderboard bars clustered nearly equal on a dark screen

Anthropic shipped Claude Opus 5 on Friday — its fourth flagship in under two months, after Mythos 5, Fable 5 and Sonnet 5. Artificial Analysis, which Anthropic engaged to evaluate the model before release, put Opus 5 at maximum effort on top of its Intelligence Index with a score of 61. Fable 5 scores 60. GPT‑5.6 Sol scores 59. Kimi K3 scores 57. Opus 4.8, the model Opus 5 replaces, scores 56.

That is five points between first place and fifth, across four labs on two continents. Read the launch coverage and you get a coronation. Read the scoreboard and you get a photo finish.

The price story has the same shape. Opus 5 bills $5 per million input tokens and $25 per million output — identical to Opus 4.8, half of Fable 5’s $10 and $50. Half the sticker. But Artificial Analysis measures what a model costs to finish the work, and there Opus 5 averages $2.03 per Intelligence Index task against Fable 5’s $2.75. That is 26% cheaper, not 50%, because a reasoning model’s bill is set by how many thinking tokens it burns, not by the rate card. Below both sit Opus 4.8 at $1.80 a task and Sonnet 5 at $1.53. The new flagship is the best value at the frontier and still the expensive way to do ordinary work.

The number in the footnotes

Buried in the same evaluation is the line nobody put in a headline. On AA‑Omniscience, which tests factual recall and deliberately rewards a model for saying it doesn’t know, Opus 5 improved accuracy by 7 points over Opus 4.8 — and its hallucination rate rose 14 points, to 50%. It got better at knowing things and worse at admitting when it doesn’t. Anthropic’s own framing is gentler, roughly 11% more accuracy for 6% more hallucination on a different base. Same direction either way: Opus 5 answers more often, and half the answers it volunteers on that benchmark are wrong.

Where it runs away with things is agentic work: 1,861 Elo on GDPval‑AA v2, 114 points clear of Fable 5, and 1,720 on AA‑Briefcase, up 146. It hits 89% on Terminal‑Bench v2.1, level with GPT‑5.6 Sol, and takes joint first on the Coding Index driven by Claude Code. Where it doesn’t: on CritPt, a frontier physics evaluation built by Argonne and UIUC researchers, it merely matches Fable 5 and trails three separate OpenAI models.

The people who used it for a week

Testers who had the model before launch weren’t running a victory lap. Dan Shipper of Every, whose team spent a week on it across coding, writing and their internal agent, called it “a hard model to love” — it argued with instructions, stopped before finishing work, and broke against skills and plugins built for previous Claudes. Their fix was to delete the scaffolding and start over, after which it performed. On Hacker News the top response was simpler confusion: if Fable 5 is the bigger model, why does Opus 5 exist? Because Fable costs twice as much and isn’t in Claude Pro. Opus 5 is now the default on Claude Max.

Our take: The capability race has compressed into a five-point band, which means the leaderboard has stopped being a buying signal. When four labs land within a rounding error of each other, the variables that decide your bill are cost per finished task, how the model behaves inside the harness you already built, and how often it makes something up. Two of those three got worse or stayed flat this launch. So ignore the index and run your own three-way: same twenty real tasks, same prompts, measure completion cost and wrong-answer rate side by side. And if you deploy Opus 5 anywhere it touches facts — research, summarization, anything customer-facing — put retrieval or a citation requirement in front of it. A model that answers more confidently is a feature in an agent loop and a liability in a knowledge base. Anthropic told you which one it optimized for.

What to watch

Advertisement

Get the day, decoded — at 7 PM ET

The Sharp Brief: AI, money, business & performance in five sharp minutes. Free.

Free bonus: subscribe today and The 2026 AI Playbook (PDF) lands with your welcome email.

Recommended by 5+ newsletters across AI, markets & business.