US Edition
Your source for latest news
OpinionEditorial

OpenAI says its newest model poses ‘critical’ cyber risk. The one outside check has been told to stay quiet.

Three frontier AI labs released cyber-capable models within the same week, each grading its own risk and setting its own safeguards. The nominal federal evaluator was told by the White House to stop talking about its work, leaving hospitals, banks and utilities to take the industry's word for it.

PN
By PressTemps NewsroomPublished Yesterday, 21:31 ET · 5 min read
OpenAI says its newest model poses ‘critical’ cyber risk. The one outside check has been told to stay quiet.
A commercial server room, illustrative file photo not depicting any of the AI labs named in this piece. Photo: BalticServers.com / Wikimedia Commons, CC BY-SA 3.0
What to know
GPT-6 Astra is the first OpenAI model to cross the "Critical" cybersecurity threshold under the company's own Preparedness Framework, scoring 100% on ExploitBench and finding two prior-unknown zero-days in testing
Google's Gemini 3.8 Flash Cyber and Anthropic's Claude Mythos 5.1 were released the same week with comparable restrictions, limiting access to vetted defenders rather than the public
The federal Center for AI Standards and Innovation (CAISI) has voluntary, non-binding pre-release testing agreements with all five major U.S. AI labs but was told by the White House in May to delete a public announcement and stop communicating publicly about its work
CAISI operates on less than $15 million a year with about 30 staff, versus a larger budget and triple the staff at Britain's AI Security Institute

On September 3, OpenAI began distributing GPT-6 Astra, the model it says is the first in its history to cross the "Critical" cybersecurity threshold defined in its own Preparedness Framework. In OpenAI's own telling, a model at that level can find previously unknown flaws in hardened, real-world systems and turn them into working exploits largely on its own. Within days, Google and Anthropic disclosed that their newest models carry comparable cyber capabilities and comparable restrictions, confirmed in a single week of overlapping announcements from all three labs. The clustering was not coincidental; it reflects how close the frontier of commercial AI has moved to capabilities once assumed to belong only to nation-states.

The numbers behind Astra's rating are specific. OpenAI reports the model scored 100 percent on ExploitBench, its internal benchmark for turning known vulnerabilities into working exploits, and discovered two previously unknown zero-day flaws during testing. The company says Astra now refuses 91.5 percent of disallowed cyber requests, up from 59 percent for its predecessor, and is being rolled out first through a limited program called Daybreak Blue rather than to the general public. Google's answer, Gemini 3.8 Flash Cyber, is being distributed only to vetted defenders through a new access program called Fairwind, which the company says already includes more than 650 government and industry partners. Anthropic, for its part, is holding back its most capable variant, Claude Mythos 5.1, restricting it to trusted-access programs for cybersecurity and life-sciences work, months after it judged an earlier Mythos preview too risky to release at all.

Three companies, three risk grades, one grader

What is missing from all three announcements is an outside check on the ratings themselves. Each company built its own capability thresholds, ran its own tests, and decided on its own safeguards before deciding what the public and paying customers would be told. There is a nominal government backstop for this. The Commerce Department's Center for AI Standards and Innovation, housed inside the National Institute of Standards and Technology, has struck voluntary pre-release evaluation agreements with every major American frontier lab, including OpenAI, Anthropic, Google DeepMind, Microsoft and xAI, and has reportedly conducted more than 40 assessments of unreleased systems. On paper, that is exactly the kind of independent verification this moment calls for.

In practice, the record is thin, and getting thinner. CAISI's own public output consists almost entirely of assessments of foreign systems, such as its published evaluation of China's DeepSeek V4 Pro, which concluded that model lagged the frontier by roughly eight months. The domestic evaluations that matter most right now, of the American labs now touching the critical-cyber threshold, are conducted behind closed doors and are not described publicly in any comparable way. Worse, when CAISI announced new testing agreements with Google DeepMind, Microsoft and xAI on May 5, the announcement was pulled from the agency's website within days at the request of the White House, which said it conflicted with a planned executive order; CAISI staff were then told to stop communicating publicly about their work and to pause interagency meetings, according to a detailed account of the agency's retreat from public view. The center now operates on less than 15 million dollars a year with roughly 30 staff, a fraction of the budget and headcount at Britain's comparable AI Security Institute.

"This feels like early COVID. There's an emergency vibe that's appropriate here," said Joshua Saxe, a former Meta AI security researcher, describing the current state of frontier-model oversight.

Who bears the cost of taking labs at their word

The people most exposed by this arrangement are not the labs themselves, who profit either way, but the institutions asked to trust their self-grading. Hospitals, utilities, banks and local governments are precisely the "critical infrastructure" defenders Google says it wants to arm through Fairwind, yet they have no independent basis for judging whether Astra's 91.5 percent refusal rate, or any other lab's stated safeguard, would hold up against a determined attacker rather than a benchmark test written by the same company. Smaller cybersecurity vendors and enterprise customers who are not invited into Daybreak Blue, Fairwind or Anthropic's trusted-access program are left further behind still, watching the most capable defensive tools get walled off from the market at the same time the underlying offensive capability, by the labs' own admission, has crossed a genuinely new threshold. National-security officials, meanwhile, are relying on a testing relationship that is voluntary in name and, per the CAISI episode, politically disposable in practice.

The industry's response to calls for binding outside review is not unreasonable on its face. Executives argue that mandatory pre-release certification, of the kind aviation regulators or the Food and Drug Administration apply to physical products, would slow releases in a race where a months-long lag can mean ceding ground to Chinese developers, whose own systems CAISI's DeepSeek evaluation shows are closing that gap faster than many assumed. They also note, accurately, that companies are already disclosing more than the law requires: Preparedness Framework write-ups, refusal-rate statistics and named access programs are all voluntary transparency measures with no legal mandate behind them. Anthropic and OpenAI can point to real internal friction, including delayed releases and withheld model variants, as evidence that self-governance has teeth. That case would be considerably stronger, however, if the one government body positioned to confirm any of it were not simultaneously being told to stop talking.

What should happen next

Voluntary self-assessment paired with a quietly muzzled evaluator is not a stable arrangement for technology that its own makers say can now discover zero-days and build cyberweapons with minimal human guidance. Congress does not need to invent a certification regime from scratch; CAISI already has the relationships, the testing protocols and forty completed assessments to build on. What it lacks is the funding, the statutory footing and the freedom to publish that would make its findings something other than an internal courtesy the White House can retract on request. Restoring CAISI's public communications, giving it a budget closer to its British counterpart's, and requiring that critical-threshold findings on domestic frontier models be disclosed in some redacted but verifiable form would not slow innovation nearly as much as labs fear, and it would give hospitals, banks and utilities something firmer than a company's own press release to rely on. Until that changes, the industry's most alarming risk ratings will remain exactly what they are today: homework graded by the students who wrote it.

More on this story

All Opinion