Sakana AI released Fugu‑Cyber on July 21 — a security‑tuned endpoint bolted onto its Fugu orchestration platform — with two headline numbers: 86.9% on CyberGym and 72.1% on CTI‑REALM. Sakana’s launch chart puts both above OpenAI’s GPT‑5.5‑Cyber and Anthropic’s Mythos Preview.
CyberGym is a UC Berkeley benchmark. It hands an agent a snapshot of a real codebase taken just before a known vulnerability was patched and asks it to reproduce the flaw — 1,507 real‑world vulnerabilities across 188 open‑source projects. It is one of the hardest agentic evaluations in circulation precisely because there is no partial credit: the attempt either triggers the expected crash or it does not.
When CyberGym’s own creators ran it at ICLR 2026, the best model‑and‑scaffold combinations cleared roughly 20%.
Sakana is claiming more than four times that. And the published benchmark graphic leaves out every field that would let anyone check: which benchmark variant, how many trials, what date, what agent scaffold, what uncertainty bounds.
Three explanations, and no way to pick one
A gap that size has a short list of causes. One: Fugu‑Cyber really is a step change, and orchestration — not model scale — is what unlocks multi‑step security work. Two: Sakana ran a different or easier subset than the authors did. Three: the scaffold is doing heavy lifting the baseline runs never had, which is a real engineering result but not the same claim as a model score.
All three are plausible. From what Sakana has published, none of them can be ruled in or out. That is the story — not that the number is wrong, but that it is unfalsifiable as printed.
The architecture matters here too. Fugu‑Cyber is not a model in the usual sense. It is a multi‑agent system that presents as a single endpoint: an orchestrator decomposes the task, routes it across a pool of specialist agents that check and challenge each other, then synthesizes one answer. Scoring that against a single model is comparing a team to an individual and reporting one number for both.
Our take: Vendor benchmarks have quietly stopped being evidence and started being marketing collateral. The tell isn’t the score — it’s the missing methodology footer. Any lab confident in a 4x result publishes the scaffold and the trial count, because reproduction is the whole point. When a chart arrives without them, treat the number as a statement of ambition, not of capability. In security specifically, buying on an unverifiable score is how you end up with a tool your team trusts more than it has earned.
What it actually costs
Access is gated: manual approval, a defensive‑use acceptable‑use policy, Token Plan only, no EU/EEA availability, no weights. Listed pricing runs $6 per million input tokens and $36 per million output, with cached input at $0.60 — each line exactly 1.2x the standard Fugu‑Ultra rate, a flat 20% premium for the cyber badge. Past 272,000 tokens of context, the rates double.
Sakana’s own guidance undercuts the leaderboard framing: the company says effective deployment still requires security specialists, verification workflows and human review. That is the honest version. It is also not what 86.9% implies to a procurement committee reading the chart.
What to watch
- Whether Sakana publishes scaffolds, trial counts and run dates — or quietly stops citing the number.
- An independent CyberGym run against Fugu‑Cyber. The benchmark is open source; the access gating is what blocks reproduction, not the eval.
- Whether rival labs answer with their own unaudited security scores. Benchmark inflation is contagious once one vendor gets away with it.
- The 20% premium holding. If orchestration is the moat, the pricing should decouple from the base model over time.
The underlying claim may well survive scrutiny. Sakana has done genuinely novel work on multi‑agent systems, and orchestration beating raw scale is a defensible thesis. But right now the largest vendor‑versus‑independent benchmark gap in recent AI is sitting in a chart with no footnotes — and the burden of proof does not belong to the reader.
