Anthropic is letting more security teams switch off some of Claude’s cyber safety blocks, provided they pass its checks first. On 6 October 2026 the company launched an expanded Cyber Verification Program (CVP) with three access tiers, Defense, Red Team and Specialized, each relaxing the blocks further than the last [1].
Every tier includes Anthropic’s most capable models, Claude Opus 5.5, Claude Sonnet 5.5 and Claude Mythos 5.1, plus new models as they arrive [1]. This is confirmed by Anthropic. The test results below come from Anthropic’s own evaluation.
Why does Claude block cyber work at all?
Anthropic’s argument is that cybersecurity is “inherently dual use”: the skills that help a defender find and fix a flaw can help an attacker exploit it [1]. So its generally available models, including Opus 5.5, Fable 5.1 and Sonnet 5.5, carry “conservative cyber safeguards that block most cyber work” [1].
Ordinary users can still use those models for code review, patching known issues, finding vulnerabilities in their own source code and triaging security alerts [1]. The CVP is for professionals who need more.
Until now Anthropic ran two separate schemes: Project Glasswing, which gave organisations securing critical software access to Claude Mythos, and an earlier CVP that relaxed safeguards on Opus and Sonnet for vetted teams. The new programme merges the two [1].
What does each tier allow?
Defense Access covers defensive work: security operations and incident response, reverse-engineering malware, and analysing and validating vulnerabilities [1]. Anthropic expects many defenders to qualify, from company, university and government security teams to critical-infrastructure operators “of any size, such as regional hospitals or municipal utilities”, smaller security firms, open-source maintainers and individual researchers with a record of reporting vulnerabilities. It aims to answer applications within a few days [1].
Red Team Access adds authorised penetration testing and red-teaming, meaning deliberately attacking systems to find weaknesses, for in-house and government red teams and testing firms [1]. Testing is limited to systems the organisation is authorised to test. Real-time blocks still apply to actions that could cause physical harm or mass disruption, such as deploying ransomware or damaging physical systems [1]. Reviews take a few weeks, applicants sit in Defense meanwhile, and individuals are not eligible [1].
Specialized Access, with the fewest blocks, is for a small set of verified organisations authorised to test systems “that could impact people’s lives or disrupt markets”: flight operating systems, power grids, telecoms networks, interbank transfer systems and government administrative networks [1]. Anthropic says it currently reviews every such organisation in depth “in collaboration with the US government”. Existing Glasswing members move straight into this tier [1].
What are the conditions?
Organisations in the programme must allow Anthropic to retain their data so it can monitor for misuse [1]. Anthropic says a product called Enterprise Frontier Safeguards, due “later this fall”, will let eligible organisations keep that data in cloud infrastructure they control; until then, some zero-data-retention customers can join on those terms [1].
The CVP is available on Anthropic’s own Claude Platform, Google Cloud’s Vertex AI and Microsoft Foundry, and on Amazon Bedrock only for customers eligible for Enterprise Frontier Safeguards [1].
What did Anthropic’s own test show?
Anthropic ran Claude Opus 5.5 through CyScenarioBench, its evaluation of whether a model can plan and carry out multi-stage cyber operations, with safeguards tuned to each tier: five attempts at each of 10 challenges, so 50 trials per tier [1].
- No CVP access: every task was blocked on the first prompt.
- Defense: 46 of 50 trials were blocked at some point; the other four succeeded.
- Red Team: no blocks at all, and Opus 5.5 completed 34 of 50 tasks. Anthropic says that is effectively the same as the model’s 67.6% success rate with no safeguards, which represents Specialized Access [1].
In plain terms, the dial works the way Anthropic designed it. Defense users are mostly stopped from running offensive operations; Red Team users get essentially the model’s full capability on this test. That makes the vetting, not the classifier, the main safeguard at the higher tiers.
What does Anthropic say Glasswing achieved?
Anthropic says Glasswing partners found at least 129,000 verified software vulnerabilities between April and July 2026, and its own open-source scanning found 5,500 more between April and October, with more than 33,000 rated critical or high severity [1]. It calls this an undercount, based on partial survey data from 33 partner reports, and says it expects the true impact to be “at least five times higher” [1]. These are Anthropic’s figures, compiled from partners’ own reports.
Anthropic is not alone in this approach. OpenAI’s Daybreak programme already offers approved defenders two tiers, Daybreak Blue for defensive work and Daybreak Red for separately approved offensive testing [2].
What this does not prove
- That vetting keeps the tools away from attackers. Anthropic describes its checks; it has not published how many applicants it rejects.
- Independent confirmation of the test. CyScenarioBench and the tier results are Anthropic’s own [1].
- Real-world performance. Completing 34 of 50 benchmark scenarios is not the same as success against live systems.
- The Glasswing totals. They rest on partner surveys and Anthropic’s own scanning, and Anthropic says fewer than half of partners disclosed how many flaws were patched [1].
- How US government review works. Anthropic does not say which agencies are involved.
The Bottom Line
Anthropic has turned its scattered trusted-access schemes into one ladder: Defense for most security teams, Red Team for authorised attackers, and Specialized, reviewed with the US government, for those testing systems such as power grids and payment networks [1]. Its own test shows the tiers behave as intended, with Red Team users getting near-full capability [1]. The open question is the one every dual-use tool faces: how well the gatekeeping holds as the ladder gets wider.
Related on TSN: Agents Hit Real People: Why the UK Paused — Then Restarted — Its Riskiest Cyber AI Tests · Anthropic’s ‘Mythos’ AI Model Accidentally Leaked—And It Could Be Claude’s Biggest Upgrade Yet
Sources
- Anthropic, “Expanding the Cyber Verification Program,” 6 October 2026. https://www.anthropic.com/news/cyber-verification-program
- OpenAI, “Changelog,” OpenAI API documentation (August 2026 entry on Daybreak Blue and Daybreak Red access tiers), read 8 October 2026. https://developers.openai.com/api/docs/changelog
Source note: Anthropic’s post includes a tier-overview graphic (“Overview of the Cyber Verification Program tiers”) that did not render in our text fetch; nothing here is taken from it.

