US Edition
Your source for latest news
OpinionArtificial Intelligence

The Companies Building AI Just Admitted They Don't Trust Themselves

A researcher's viral resignation from Anthropic, corroborated on the record by the company's own safety staff, exposed how thin voluntary AI guardrails have become — and handed Congress a rare opening it should not waste.

PN
By PressTemps NewsroomPublished Yesterday, 17:50 ET · 6 min read
The Companies Building AI Just Admitted They Don't Trust Themselves
The U.S. Capitol, where lawmakers are now weighing competing bills on AI safety following a viral resignation from an Anthropic researcher. Photo: 颐园居 / Wikimedia Commons, CC BY-SA 4.0.
What to know
Anthropic researcher Jacob Coxon resigned Sept. 9 warning that AI labs are "racing straight to self-improving superintelligence"; his thread drew tens of millions of views
Anthropic's own alignment science lead, Evan Hubinger, publicly agreed and put the odds of AI-caused extinction within a decade at over 10 percent
Anthropic loosened its Responsible Scaling Policy in February 2026, replacing hard-pause triggers with "nonbinding but publicly-declared targets," even as internal risk warnings intensified
Congress is now weighing two competing responses — a Sanders-Casar bill to ban superintelligent AI outright and a more incremental bipartisan Klobuchar-Cruz-Thune bill — and the piece argues for binding disclosure and liability over either extreme

On Tuesday, a 27-year-old researcher named Jacob Coxon resigned from Anthropic and posted a seven-part thread on X that, within days, had been viewed tens of millions of times. His message was blunt: after three years doing pretraining research at both OpenAI and Anthropic, he had concluded that "neither company is acting responsibly," that both were "racing straight to self-improving superintelligence and gambling with our lives," and that the people building the technology "earnestly believe that it could kill us all by the end of the decade." What made the moment different from the AI industry's previous safety scares was what happened next: Anthropic's own staff got up and agreed with him.

Evan Hubinger, who leads alignment science at Anthropic, wrote that Coxon was "correct" and that he personally estimates the odds of AI causing human extinction within the next decade at more than 10 percent.

"Jacob is correct here — we really do earnestly believe AI could kill all humans."

Samuel Marks, who works on scalable oversight at the company, said much the same in a personal capacity, adding that more senior Anthropic employees tend to be more, not less, alarmed. This is the detail that should unsettle Washington more than Coxon's resignation itself. It is one thing for a departing employee to sound an alarm; skeptics can wave that away as a disgruntled ex-staffer chasing attention. It is another for the scientists still cashing the company's paychecks to stand up in public and say, in effect, that he was right.

The safeguards got softer as the warnings got louder

The timing compounds the irony. Anthropic revised its own Responsible Scaling Policy this year, and the update, effective in February, replaced the earlier framework's hard, trigger-based commitments with what the company itself now describes, in its explanation of version 3.0 of the policy, as "nonbinding but publicly-declared targets" laid out in a Frontier Safety Roadmap, backed by periodic risk reports rather than an automatic pause. In other words, at almost the same moment its own alignment lead was telling the public he takes double-digit-percentage extinction odds seriously, the company was loosening the one voluntary mechanism designed to stop deployment if things went wrong.

None of this is happening on the fringe of AI research. A 2022 survey of machine-learning researchers who had published at top AI conferences found a median estimate of 5 percent that advanced AI's long-run effect on humanity would be "extremely bad, e.g., human extinction," with 15 percent of respondents putting the odds at 25 percent or higher — figures that have held roughly steady since the same survey was run in 2016. Coxon and Hubinger are not describing some novel fear invented for a viral thread; they are restating, with company letterhead attached, an assessment that a meaningful slice of the field has held for a decade.

Why this warning, and not the last one

AI researchers have quit and sounded alarms before. When OpenAI's Jan Leike resigned in 2024 over safety concerns, his statement drew a fraction of the attention Coxon's has — by one comparison, roughly 6 million views versus what some outlets are now tallying as upward of 100 million for Coxon's thread. What changed is context, not just content. Months of public backlash over data centers' land, water and electricity demands had already primed Americans to distrust the industry's promises: polling cited in the days after Coxon's post found 64 percent of Americans opposed to rapid data-center expansion and 77 percent worried about the effect on their electricity bills. Into that atmosphere, a credentialed insider's warning, corroborated by name from inside the company he was criticizing, landed differently than a similar warning would have two years ago.

It landed differently in Washington, too. Within a week, Senator Bernie Sanders and Representative Greg Casar had unveiled legislation to permanently ban the development of superintelligent AI systems and impose a temporary pause on advanced model development until a new federal regulator writes safety rules, with penalties for individuals of up to 20 years in prison — the range reserved for illegal nuclear-weapons development — and dissolution for violating companies. Separately, a more incremental, bipartisan effort from Senators Amy Klobuchar, Ted Cruz and Majority Leader John Thune, aimed at "commonsense guardrails" against catastrophic biological and nuclear misuse of AI, is reportedly gaining momentum and could be introduced within days.

Who actually bears the risk

It is worth being precise about who is exposed here, because the debate too easily collapses into an abstraction. It is not only some hypothetical future public confronting a rogue superintelligence. It is AI-lab employees who must now decide, as Coxon did, whether raising internal concerns is enough, or whether the only real lever left is a public resignation — which is why Anthropic's own policy update also included an expanded whistleblower and anti-retaliation policy, updated in February that Coxon's resignation will now test in public. It is investors and taxpayers exposed to an AI sector now carrying more than $1.5 trillion in corporate borrowing. And it is ordinary ratepayers already feeling the electricity and water costs of the data-center buildout that all of this frontier-model competition depends on. The externalities of the AI race are not evenly distributed between the labs racing each other and everyone else.

There are, to be fair, real grounds for caution about overcorrecting. Hubinger himself clarified that his 10-percent estimate concerns future, more capable systems built through recursive self-improvement, not today's models, which Anthropic's own risk assessments describe as low-risk. Critics of the Sanders-Casar approach — including inside the industry — argue that a blanket ban enforced with prison sentences invites definitional chaos over what counts as "superintelligent," risks pushing frontier research to jurisdictions with no safety culture at all, and could not plausibly pass a Republican-controlled Congress that has spent the past year deregulating AI at the state level. OpenAI's head of global affairs, Chris Lehane, put the more cautious industry position plainly, arguing lawmakers need "a bias toward meaningful action over policy perfection" rather than a maximalist bill that goes nowhere.

What should happen next

That argument for pragmatism over purity is worth taking seriously — but it cuts against relying on voluntary corporate policy just as forcefully as it cuts against Sanders' 20-year prison terms. Anthropic's own scaling policy has now been revised three times in three years, each time by the company that wrote it, judged against standards the company set for itself, with compliance reported on a schedule the company controls. That is not oversight; it is the industry grading its own homework, and even its safety researchers are publicly saying they are not confident in the grade. The more durable path is the one the bipartisan Senate effort gestures toward but has not yet delivered: mandatory, independently verified risk disclosure; binding pause authority that does not depend on a company's internal roadmap; and real liability when a lab's own safety commitments turn out to be aspirational rather than operative. Congress has let AI-safety alarms fade into the news cycle before. This one came with the industry's own signature attached. That should be harder to ignore, and it should not take another viral thread to make it stick.

More on this story

All Opinion