Claude AI system interface displaying mathematical equations and computational results related to Riemann zeta function analysis

Anthropic Researcher Issues Warning Over AI Labs Racing to Build Self-Improving Systems

Anthropic researcher Jacob Coxon resigned this week, warning that leading AI labs including Anthropic and OpenAI are “gambling with our lives” by racing to build self-improving AI systems he described as out of control.

Coxon, who specialized in training large models by having them consume vast amounts of data, told the Wall Street Journal that he is leaving the AI industry entirely, not just Anthropic. Before joining Anthropic, his profile lists him as a member of technical staff at OpenAI since 2020, focused on AI safety and model interpretability.

The timing is pointed, Anthropic is reportedly in a critical IPO preparation period, according to TradingKey, making a public safety resignation from inside the company an unusually sensitive event for a lab about to court public shareholders.

A Named Anthropic Researcher Departure, Not An Anonymous Leak

Coxon’s exit differs from the usual pattern of anonymous safety complaints inside frontier labs. He put his name and résumé behind the warning, telling reporters he believes AI could pose existential risk to humanity within the decade if the current competitive dynamic continues unchecked.

As an Anthropic researcher with direct exposure to how large models are built and trained, his credentialed insider account carries a different weight than outside commentary, and raises questions worth tracing back to primary sources, including model cards and published research, rather than taking at face value from any announcement.

His argument centers on self-improving models, the idea that once an AI system can meaningfully improve its own training or successor systems, the pace of capability gains could outstrip any lab’s ability to test or contain them.

Whether that threshold is near or distant is precisely the kind of empirical question that benchmark numbers alone do not settle, they measure performance on defined tasks, not the open-ended dynamics Coxon is describing.

Echoes From Outside The Labs

The concerns raised by Coxon arrive alongside separate warnings from mathematician Terence Tao, who has spent time examining what current AI systems can and cannot reliably do, according to Gary Marcus. Neither Anthropic nor OpenAI has issued a detailed public response to Coxon’s specific claims.

The convergence of an insider account and a mathematician of Tao’s stature raising independent concerns from outside suggests the debate is widening beyond its usual participants, right as Anthropic’s commercial ambitions face their biggest public test yet.

Read Next: Anthropic’s IPO Push Faces New Scrutiny Over AI Safety Claims

Similar Posts