Anthropic ran a protein design campaign against 15 targets, handed the results to two outside labs, and got the wet-lab data back. Claude designed working binders against 14 of them. Between 22% and 35% of individual designs actually bound, depending on the setup. The industry norm for de novo binder campaigns is 10–15%.
The number that matters more is who did the work. Anthropic wrote a prompt, granted network approvals, and left. No scientist steered the campaign, chose the binding sites, or picked which structure-design model to run. Claude selected the epitopes, orchestrated several open-source structure, sequence and co-folding models, ran the in-silico optimisation cycles, and screened for candidates that would express and stay soluble — the multi-week loop a protein engineer normally grinds through per target. Adaptyv Bio and Twist Bioscience then independently made and tested the designs.
Out of 1,320 designs, 354 confirmed binders came back. For scale, the two largest public collections of de novo binders together hold roughly 770 binders from about 5,700 designs across 40 targets. One autonomous campaign moved the public corpus by something close to half.
The competitive test
Benchmark hit rates are easy to inflate when the targets are well-studied and the answers sit in training data. Anthropic pulled two targets — 15-PGDH and GDF-8 — from Adaptyv Bio's most recent open competitions, and required Claude to verify its designs were original.
On RBX1, where Adaptyv had run a public competition, Claude in single-target mode hit 40%. The human field in that competition hit 3.7% across 245 entries. Claude's top design bound more tightly than the winner.
The mode mattered. Running all targets at once in a single 48-hour session, Mythos Preview hit 26.7% and Opus 4.8 hit 22.6%. Given one target per session across parallel 24-hour runs — closer to how a human engineer actually works — Mythos Preview reached 35.1%.
Our take: The headline is 14 of 15, but the finding underneath is about orchestration, not chemistry. Claude did not invent a better protein-design model. It ran the field's existing open-source tools better and faster than the specialists who built them, for 24 to 48 hours, unsupervised. That is a general claim about agentic work, and it does not stay inside biology.
Where it broke
Anthropic published the failures, which is the part worth trusting. Against maltose-binding protein, none of 90 designs was confirmed — a large, flexible surface with nothing to grip. Against BBF-14, a synthetic beta-barrel built specifically to have no natural precedent, Claude managed three binders with only modest affinity.
There is also an unexplained inversion: on TNFα, the target behind Humira and one that expert groups have repeatedly failed on, the older Opus 4.8 succeeded where the more capable Mythos Preview did not. Anthropic says it does not know why. Capability on these tasks is lumpy, not a clean ladder.
The chemistry half
The second experiment used Opus 5, a generally available model. Given a contract lab's raw NMR and LC-MS instrument files — proprietary binary formats meant to be opened only in vendor software — and a two-sentence prompt, it returned finished analyses in 23 and 19 minutes running in parallel. Purity came back at 96.4% against the lab's 96.33%. Hydrogen counts landed within 0.08 ¹H.
It also reverse-engineered the undocumented LC-MS format, then proved it had read the file correctly by reproducing the instrument's own totals across all 2,664 scans before analysing anything. And it proposed the same heavy-water follow-up experiment the lab had independently run — then caught itself overstating the first result and corrected it. The lab's finished report for that sample arrived four days after the first spectrum.
What to watch
- The access program. Protein design stays blocked in Anthropic's most capable model on dual-use grounds. The gated scientist program it says is coming is where this becomes usable rather than demonstrable.
- Replication. Anthropic says it intends further characterisation to confirm the hit rates. Two external labs is good; independent groups reproducing the campaign from the published prompts is better. The prompts and data are public.
- The gap to a drug. Minibinders are not a standard therapeutic modality, and a high-affinity binder is the first step of many. Nothing here shortens a clinical trial.
- The transferable bit. A model that can run other people's specialist software unattended for 48 hours is the story. Watch for the same pattern in materials, semiconductors, and anywhere a workflow is bottlenecked on expert orchestration rather than expert insight.
Anthropic's own framing is careful: the bottlenecks in drug development are mostly policy and operations, not raw scientific capability. That is true, and it is also the part no model fixes. But the weeks-per-target orchestration tax was real, and this result says it is negotiable.
