US Edition
Your source for latest news
TechnologyCybersecurity

Anthropic Widens Access to AI Hacking Tools, Citing 129,000 Vulnerabilities Found

The company folded its Project Glasswing critical-infrastructure program into a three-tier system that loosens restrictions on its models for vetted security teams, even as independent researchers question how many of its claimed findings have actually been verified and patched.

PT
By PressTemps Technology DeskPublished Yesterday, 21:35 ET · 7 min read
Anthropic Widens Access to AI Hacking Tools, Citing 129,000 Vulnerabilities Found
Jack Clark, Anthropic's co-founder and head of policy, speaking on AI at the Schwarzman Centre, Oxford. (Photo via Wikimedia Commons, CC0 1.0 — uploader Mvolz)
What to know
Anthropic merged Project Glasswing and its Cyber Verification Program into three access tiers — Defense, Red Team and Specialized — covering Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1.
The company says partner organizations identified at least 129,000 verified software vulnerabilities between April and July, with more than 33,000 rated critical or high severity.
An independent VulnCheck analysis found only about 2,700 of roughly 26,000 claimed findings have reached Anthropic's public disclosure ledger, with just 70 to 82 assigned CVE numbers so far.
Organizations in the program must let Anthropic retain usage data to monitor for misuse until a new data-control service, Enterprise Frontier Safeguards, arrives later this fall.

Anthropic said Tuesday that it is restructuring and widening the program that governs how security teams can use its most capable artificial-intelligence models for offensive cyber work, folding its six-month-old Project Glasswing initiative into a single framework with three levels of access. The change, described in a post on the company's website, is the clearest sign yet that Anthropic intends to treat AI-assisted vulnerability hunting as a mainstream defensive tool rather than an experiment confined to a handful of vetted partners.

The expanded Cyber Verification Program, or CVP, replaces a looser arrangement under which Project Glasswing gave a limited set of critical-infrastructure operators access to an unreleased model called Claude Mythos, while a separate and smaller verification track gave vetted security teams reduced guardrails on Anthropic's public models. The two are now combined into tiers named Defense Access, Red Team Access and Specialized Access, each covering Claude Opus 5.5, Claude Sonnet 5.5 and Claude Mythos 5.1.

A Tiered System Replaces Ad Hoc Access

Under the new structure, Anthropic's announcement describes Defense Access as the broadest and fastest-moving tier, open to companies, nonprofits, universities, government agencies and even individual researchers with a record of reporting vulnerabilities. It covers security-operations work, incident response, malware reverse-engineering and vulnerability analysis, and the company says it aims to approve applications within a few days.

Red Team Access adds authorized penetration testing and red-teaming, but is limited to organizations testing systems they are explicitly permitted to test; Anthropic says review takes a few weeks, during which applicants are provisionally enrolled in Defense Access. Specialized Access, the least restricted tier, is reserved for a small number of verified organizations cleared to test systems such as flight-operating software, power grids, telecommunications networks and interbank transfer infrastructure, with each applicant reviewed jointly with the U.S. government. Existing Project Glasswing members move automatically into Specialized Access for the models they already use.

Enrollment comes with a condition: organizations in the program must agree to let Anthropic retain usage data so it can monitor for misuse. The company says a forthcoming service called Enterprise Frontier Safeguards, expected later this fall, will let eligible organizations keep that data in infrastructure they control rather than on Anthropic's servers. The tiers are available through Anthropic's own platform, Google Cloud's Vertex AI and Microsoft Foundry; Amazon Web Services' Bedrock service will offer the program only to customers that qualify for the forthcoming safeguards service.

The Numbers Behind the Expansion

Anthropic disclosed new figures on how much vulnerability-hunting activity the program has produced. Partners in Project Glasswing identified at least 129,000 verified software vulnerabilities between April and July of this year, the company said, and its own scanning of open-source code turned up another 5,500 between April and October. More than 33,000 of the combined total have been rated critical or high severity. Anthropic cautioned that the figure is likely an undercount, since it is drawn from survey responses submitted by only a subset of partners, and said it expects the real impact to be at least five times higher once more partners report their patch data.

To justify loosening the model's restrictions for vetted users, Anthropic published results from an internal benchmark called CyScenarioBench, which ran Claude Opus 5.5 through ten attack-style challenges, five attempts each, at every tier. With no program access, all 50 attempts were blocked on the first prompt. Under Defense Access, 46 of 50 attempts were blocked at some point, with four succeeding. Under Red Team Access, none were blocked and the model completed 34 of 50 tasks — a completion rate Anthropic said matched the 67.6 percent success rate it measured when the model was tested with no safeguards at all.

Independent researchers have offered a more skeptical count of the program's real-world impact. A VulnCheck analysis published in September found that of roughly 26,000 findings Anthropic has claimed through Glasswing, only about 2,700 have reached its public disclosure ledger, and of those, just 70 to 82 have been assigned CVE identifiers so far. VulnCheck's Patrick Garrity, who conducted the analysis, wrote that "the receipts are starting to trickle in, they just don't reconcile," pointing to a wide and still-unexplained gap between the volume of findings Anthropic claims and the much smaller number that have been independently verified, patched and published.

From a Narrow Pilot to a Formal Policy

Anthropic launched Project Glasswing in early April with roughly a dozen founding partners, including Amazon Web Services, Cisco, Google, Microsoft, Nvidia, JPMorganChase and the Linux Foundation, giving them gated access to Mythos Preview, a model the company said was too capable in offensive cyber tasks to release broadly. Anthropic pledged up to $100 million in usage credits and $4 million in donations to open-source security groups as part of the effort, and said the model had already surfaced zero-day flaws including a decades-old weakness in OpenBSD and a long-standing bug in the FFmpeg video library. The program expanded in June to roughly 150 additional organizations across more than 15 countries, adding sectors such as power, water and healthcare that were absent from the initial roster.

That expansion ran alongside a separate, smaller Cyber Verification Program that gave vetted security teams reduced restrictions on Anthropic's publicly available models. This week's announcement merges the two, a step Anthropic frames as part of a broader argument it has been making since Glasswing's launch: that cybersecurity tools are inherently dual-use, and that the same AI capability that helps a defender find a flaw before an attacker does will eventually be available from other AI labs regardless of what Anthropic does. The backdrop is a security industry already grappling with AI's uneven effect on code quality — a separate report from the security firm Veracode found that roughly 44 percent of AI-generated coding tasks introduced a risky vulnerability, with an average security pass rate of 56 percent, little changed from a year earlier.

Security Researchers Offer a Mixed Verdict

Reaction from outside analysts has been measured. Sakshi Grover, an analyst at IDC quoted by CSO Online, noted that Anthropic's benchmark was run by the company itself rather than an independent evaluator.

"A vendor-run test is not the same as an independent one. Whether this kind of evaluation is legitimate depends less on the technique and more on the authorization, scope and execution conditions," Grover said, adding that enterprises should pair any reduced model restrictions with controls outside the model itself, such as distinct identities for AI agents, short-lived task-specific privileges and independent checks on both targets and actions.

Deepika Giri, also of IDC, said the tiered approach "works only if vetting is real and access is continuously rechecked," since an AI agent's behavior can change over time, and recommended that organizations treat agents with elevated cyber permissions the way they would treat a privileged employee — with short-lived access, full observability and human sign-off on high-impact actions. A third researcher, Vibhum Dubey, told the outlet that organizations deploying the tools still need to settle who is accountable when an autonomous system takes a consequential action.

The Hacker News, which first flagged the scale of the vulnerability figures, noted separately that an analysis of 300 vulnerabilities attributed to Anthropic or Glasswing partners found only two had been confirmed as exploited in the wild: a SQL-injection flaw in the Ghost content-management system and a session-forgery bug in the Rejetto HTTP File Server, both now patched. That roughly half-percent exploitation rate is consistent with VulnCheck's broader finding that AI-discovered vulnerabilities, so far, are being weaponized by attackers at about the same rate as vulnerabilities found through conventional means — undercutting, for now, fears that AI-assisted bug-hunting is primarily arming attackers rather than defenders.

Anthropic has not said how many organizations have applied for the new tiers or when it expects to publish updated figures. The company said it will continue to adjust its access classifiers "over time" as it gathers more data on how each tier is used, and that the Enterprise Frontier Safeguards service meant to let partners keep their own data remains on track for later this fall. For now, the program's success will likely be judged less by the raw count of vulnerabilities Anthropic can claim and more by how many of them are independently confirmed, patched and kept out of the hands of attackers — the gap that researchers like Garrity say the company has yet to close.

More on this story

All Technology