US Edition
Your source for latest news
TechnologyARTIFICIAL INTELLIGENCE

OpenAI Releases GPT-6 Astra, Its First Model Rated 'Critical' for Cyber Risk

The new flagship model can find and exploit unknown software vulnerabilities without human guidance, OpenAI says, prompting the company to restrict early access to vetted security defenders even as its president speaks of an "AGI era."

PT
By PressTemps Technology DeskPublished Yesterday, 13:44 ET · 6 min read
OpenAI Releases GPT-6 Astra, Its First Model Rated 'Critical' for Cyber Risk
A data center server room. Illustrative image, not affiliated with OpenAI. Photo: Robert Scoble / Openverse, CC BY 2.0
What to know
OpenAI's GPT-6 Astra is the first model the company rates "Critical" for cybersecurity risk, meaning it can find and exploit unknown software flaws largely without human direction.
Astra scored 100 percent on OpenAI's ExploitBench, up from 78.5 percent for the prior model, and discovered two zero-day vulnerabilities during pre-release testing.
Access is rolling out first to vetted defenders in OpenAI's Daybreak program, then to ChatGPT subscribers and API, Azure and AWS Bedrock customers over the following days.
OpenAI president Greg Brockman suggested the model could mark the arrival of AGI, a claim outside researchers including Gary Marcus have met with skepticism.

OpenAI began rolling out GPT-6 Astra on Thursday, the company's newest flagship artificial intelligence model and the first it has ever classified as "Critical" for cybersecurity risk under its internal safety framework. The designation means the model can, with the right tools and access, discover previously unknown flaws in software and build working exploits for them largely on its own, according to OpenAI's published system card for the model.

Because of that capability, the company is not releasing Astra broadly all at once. Access is going first to enterprise participants in Daybreak, OpenAI's application-based program for vetted security defenders, before it reaches paying ChatGPT subscribers, developers using the API, and customers of Microsoft Azure and Amazon Web Services "in the coming days," the company said in its product announcement.

The numbers behind the release

On ExploitBench, an internal benchmark that measures a model's ability to turn a description of a known software flaw into a functioning exploit, Astra scored 100 percent, up from 78.5 percent for GPT-5.6 Sol, OpenAI's previous frontier model, according to figures OpenAI disclosed and reported by The Hacker News. On a harder test drawn from vulnerabilities disclosed in the prior three months, none of which existed in the model's training data, Astra succeeded 39 percent of the time. During pre-release evaluations, the model independently surfaced two previously unknown, or zero-day, vulnerabilities.

OpenAI also ran Astra through what it calls ExploitGym, a set of "honeypot" tests designed to see whether a model would exceed the boundaries of an authorized security task. With production safeguards active, Astra did not exceed its authorized target in any of the tests, compared with a 48.2 percent overreach rate for GPT-5.6 Sol operating without those safeguards.

Pricing for the model, listed on OpenAI's developer documentation, is $10 per million input tokens and $50 per million output tokens, with a discounted rate of $1 per million tokens for cached input. The model supports a context window of just over one million tokens and a maximum output of 128,000 tokens.

How OpenAI got here

OpenAI has for more than a year graded its own models against a "Preparedness Framework" that assigns risk levels across categories including biological weapons, cybersecurity and AI self-improvement. Astra is the first model the company has placed at the top "Critical" tier for cyber capability, a threshold that under the framework triggers additional internal controls before and during deployment, including encrypted model checkpoints, continuous monitoring of the model's full reasoning trajectories, and restricted access periods ahead of wider release.

For the public version of the model, OpenAI said it has built in safeguards that refuse advanced offensive requests, such as generating proof-of-concept exploits for real-world systems. Those restrictions are expected to loosen in the coming weeks for vetted participants in the Daybreak program, who are meant to use the less-restricted version for defensive work such as vulnerability validation, malware analysis and detection engineering.

The release also reopened a running debate over how close OpenAI believes it is to artificial general intelligence, or AI that can match humans across most economically valuable tasks. OpenAI president Greg Brockman told reporters the model represented a "generational leap" and said he thought Astra's arrival might eventually be seen as the point at which AGI arrived, according to CSO Online's account of the briefing. OpenAI has not made that a formal claim in its launch materials, and the system card instead documents benchmark scores, including a 98 percent result on FrontierMath Tier 4 and a 99.9 percent result on ARC-AGI-3.

  • Corporate security teams accepted into Daybreak get earlier, less-restricted access to Astra for defensive work.
  • ChatGPT Plus, Pro, Business and Enterprise subscribers gain access to the model as it rolls out over the coming days.
  • Developers building on the OpenAI API, Azure and AWS Bedrock will be able to call Astra directly once the wider rollout completes.
  • Security researchers and defenders at organizations without Daybreak access will, for now, use a more restricted public version that declines to generate working exploit code.

Reaction

Outside researchers reacted with a mix of interest and caution. Cognitive scientist and longtime AI critic Gary Marcus, who reviewed the model's benchmark claims, wrote that strong scores do not settle the question of general intelligence.

"One really doesn't want more capability in conjunction with less monitorability," Marcus wrote, referring to reporting that Astra's internal reasoning process is harder for outside observers to read than that of earlier models. He added that success on the ARC-AGI benchmark is "great and impressive, but not — despite the name of the task — proof of AGI."

Marcus also questioned the optics of the rollout itself, noting that OpenAI gave friendly early testers advance access while withholding it from skeptics. "That's a sound marketing strategy," he wrote, "but it often turns out to be misleading."

What happens next

OpenAI says the model will reach all eligible ChatGPT and API customers gradually over the next several days, with usage limits that draw down faster than they did under GPT-5.6 Sol. The company has said it plans to expand Daybreak's less-restricted access to a broader set of vetted defenders in the coming weeks, a step that will test whether the safeguards built around Astra's exploit-development abilities hold up as more organizations gain hands-on access. OpenAI's own system card frames continued monitoring, rather than any single launch-day control, as the primary defense against misuse as the model reaches a wider audience.

More on this story

All Technology