Anthropic Researcher Warns of Dangerous AI Risk, Cites 10% Extinction Odds

An Anthropic researcher named Jacob Coxon publicly resigned this week, warning that the industry’s rush toward self-improving AI has become “out-of-control”, prompting alignment lead Evan Hubinger to put the odds of AI-caused human extinction within a decade at above 10%.

Coxon’s exit, first detailed by the Wall Street Journal, frames his departure not as a rejection of the company’s mission but as a warning that the broader race among AI labs to build increasingly autonomous, self-improving systems has outpaced anyone’s ability to control the outcome.

Also Read: Anthropic Researcher Issues Warning Over AI Labs Racing to Build Self-Improving Systems

He has reportedly told colleagues that companies are “playing with our lives” by continuing to scale systems whose behavior even their own safety teams cannot fully predict.

Hubinger’s response is the more striking data point. Rather than distancing the lab from Coxon’s alarm, he affirmed it, putting a specific probability, greater than 10%, on AI causing human extinction within ten years, according to CNBC. That figure comes from a safety researcher inside a company whose entire public identity rests on being the cautious alternative to OpenAI and other labs.

Internal Dissent At A Safety-First Lab

Anthropic was founded in 2021 by former OpenAI staff, including Dario Amodei and Daniela Amodei, explicitly to build AI more carefully than rivals. Coxon’s resignation is notable precisely because it comes from inside that organization rather than from an outside critic.

Policy Response And Open Questions

Neither the company nor Hubinger has announced any policy change following the resignation, and its public materials still describe the mission as building “reliable, interpretable, and steerable AI systems,” per its own site.

Hubinger’s stated probability raises a harder question worth tracing to the safety team’s own model cards and internal benchmarks, what exactly are those tools measuring about system controllability, and do the numbers there support or complicate his estimate? Whether that figure shifts release timelines or triggers further departures remains unclear, and worth watching.

Read Next: Inside AI Labs’ Race to Self-Improving Systems

Similar Posts