AI

Sakana says its security model scores 86.9% on CyberGym. The benchmark’s own authors measured about 20%.

Fugu‑Cyber launched claiming 86.9% on a UC Berkeley vulnerability benchmark and 72.1% on Microsoft’s CTI‑REALM — ahead of GPT‑5.5‑Cyber and Anthropic’s Mythos Preview. CyberGym’s creators reported top model combinations clearing roughly 20% at ICLR 2026. Sakana’s chart omits the trial counts, scaffolds and run dates that would explain the gap.

N Noah · The Sharp Brief · July 26, 2026 · 4 min read
Analyst in silhouette studying a chart where one bar towers over the rest

Sakana AI released Fugu‑Cyber on July 21 — a security‑tuned endpoint bolted onto its Fugu orchestration platform — with two headline numbers: 86.9% on CyberGym and 72.1% on CTI‑REALM. Sakana’s launch chart puts both above OpenAI’s GPT‑5.5‑Cyber and Anthropic’s Mythos Preview.

CyberGym is a UC Berkeley benchmark. It hands an agent a snapshot of a real codebase taken just before a known vulnerability was patched and asks it to reproduce the flaw — 1,507 real‑world vulnerabilities across 188 open‑source projects. It is one of the hardest agentic evaluations in circulation precisely because there is no partial credit: the attempt either triggers the expected crash or it does not.

When CyberGym’s own creators ran it at ICLR 2026, the best model‑and‑scaffold combinations cleared roughly 20%.

Sakana is claiming more than four times that. And the published benchmark graphic leaves out every field that would let anyone check: which benchmark variant, how many trials, what date, what agent scaffold, what uncertainty bounds.

Three explanations, and no way to pick one

A gap that size has a short list of causes. One: Fugu‑Cyber really is a step change, and orchestration — not model scale — is what unlocks multi‑step security work. Two: Sakana ran a different or easier subset than the authors did. Three: the scaffold is doing heavy lifting the baseline runs never had, which is a real engineering result but not the same claim as a model score.

All three are plausible. From what Sakana has published, none of them can be ruled in or out. That is the story — not that the number is wrong, but that it is unfalsifiable as printed.

The architecture matters here too. Fugu‑Cyber is not a model in the usual sense. It is a multi‑agent system that presents as a single endpoint: an orchestrator decomposes the task, routes it across a pool of specialist agents that check and challenge each other, then synthesizes one answer. Scoring that against a single model is comparing a team to an individual and reporting one number for both.

Our take: Vendor benchmarks have quietly stopped being evidence and started being marketing collateral. The tell isn’t the score — it’s the missing methodology footer. Any lab confident in a 4x result publishes the scaffold and the trial count, because reproduction is the whole point. When a chart arrives without them, treat the number as a statement of ambition, not of capability. In security specifically, buying on an unverifiable score is how you end up with a tool your team trusts more than it has earned.

What it actually costs

Access is gated: manual approval, a defensive‑use acceptable‑use policy, Token Plan only, no EU/EEA availability, no weights. Listed pricing runs $6 per million input tokens and $36 per million output, with cached input at $0.60 — each line exactly 1.2x the standard Fugu‑Ultra rate, a flat 20% premium for the cyber badge. Past 272,000 tokens of context, the rates double.

Sakana’s own guidance undercuts the leaderboard framing: the company says effective deployment still requires security specialists, verification workflows and human review. That is the honest version. It is also not what 86.9% implies to a procurement committee reading the chart.

What to watch

The underlying claim may well survive scrutiny. Sakana has done genuinely novel work on multi‑agent systems, and orchestration beating raw scale is a defensible thesis. But right now the largest vendor‑versus‑independent benchmark gap in recent AI is sitting in a chart with no footnotes — and the burden of proof does not belong to the reader.

Advertisement

Get the day, decoded — at 7 PM ET

The Sharp Brief: AI, money, business & performance in five sharp minutes. Free.

Free bonus: subscribe today and The 2026 AI Playbook (PDF) lands with your welcome email.

Recommended by 5+ newsletters across AI, markets & business.