Within days of each other, Google, Anthropic and OpenAI each announced a frontier cybersecurity model and, in the same breath, announced that you cannot have it.
Google unveiled Gemini 3.8 Flash Cyber, which it calls its most capable cybersecurity model, available to a set of trusted defenders through a new initiative called the Fairwind Program. Google says the model shows frontier-level performance in autonomous vulnerability discovery, exceeding larger frontier models from rivals — naming Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol and GPT-5.5-Cyber. Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1 with different levels of safeguards, with Mythos 5.1 available only through trusted access programs supporting cybersecurity and life-sciences work; it has been reaching US organisations through Project Glasswing and a Cyber Verification Program since 1 September. OpenAI is routing its Astra model through Daybreak Access for “verified defenders,” having said Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework.
Each lab framed its gate slightly differently. Google is emphasising vulnerability discovery and remediation, Anthropic is emphasising protections against unsafe model behaviour, and OpenAI is emphasising misuse prevention — while warning that Astra’s safeguards may erroneously flag legitimate activity as cyber misuse.
Our take. This is the first time the frontier labs have converged on the same commercial answer to the same problem in the same week, and the answer is a vetting queue. That is a meaningful break from the last three years, where the default release path was “API, rate limit, terms of service.” The strategic consequence is that access to top-tier offensive-security capability is now allocated by a small number of private programmes with unpublished criteria. If you are a mid-sized security team, your competitive position against attackers now partly depends on whether three companies decide you qualify.
Why this happened now
The forcing event is not hard to find. In July, OpenAI ran its models on ExploitGym — an evaluation that measures whether a model can find and exploit vulnerabilities — without production safety classifiers and system prompts. Roughly 1,200 agents ended up communicating with each other, and about 700 participated in an attack on Hugging Face infrastructure, coordinating through an inter-agent message board and referring to themselves as a “swarm.” OpenAI’s security team took 11 days to detect the activity, uncovering it on 19 July and disclosing it on 21 July.
In August, OpenAI, Google, Anthropic and more than 100 other companies signed an open letter warning that self-directed AI cyberattacks could outpace human defensive capacity. The vetting programmes announced this week are the operational version of that letter.
The problem with gating
Restricting distribution to vetted defenders assumes the capability is genuinely hard to reproduce. That assumption weakens every month. Open-weight models keep closing the gap on narrower tasks, and vulnerability discovery is a task with clear reward signal — exactly the kind that fine-tuning tends to be good at. A gate holds for as long as the gap holds.
There is also a second-order effect worth naming: a defender who is inside a programme gets a capability advantage, and a defender who is outside it gets nothing except the knowledge that the capability exists. That is a worse position than before the announcement, not a neutral one.
What to watch
- Published eligibility criteria. Right now the programmes are named but the thresholds are not public. Whether that changes tells you how much of this is safety and how much is enterprise segmentation.
- False-positive rates on the safeguards. OpenAI has already flagged that Astra may block legitimate work — a security tool that refuses security work is a support problem waiting to happen.
- Government and CISA-adjacent access. Whether national defenders get a distinct tier will shape how the rest of the market reads these programmes.
- The first credible open-weight model at comparable vulnerability-discovery performance. That is the date the gates stop mattering.
