AI Incidents
Back to published register
Risk signalInsider warningNot an incident

Former Anthropic safety researcher Joe Benton warns about loss of control and calls for transparency

Joe Benton said he left Anthropic because, in his view, leading AI companies underinvest in safety and do not give the public enough visibility into loss-of-control risks, safety incidents, or progress toward self-improving systems. He is joining METR to support independent evaluations. His statements are an expert warning, not empirical evidence that loss of control will occur.

Published
Sep 11, 2026
Signal type
Insider warning
Confidence
99%
Organization
Anthropic, METR
Sources
3
Last reviewed
Sep 14, 2026

Assessment

What this signal changes

Benton says he led a scalable-oversight team at Anthropic and describes a structural conflict between the race for more capable systems and adequate safety investment. He calls for reporting of incidents and near misses, minimum standards, and independent assessments. His original essay, personal website, and independent Associated Press reporting establish his role, move to METR, and core statements.

Evidence boundary

What the evidence does not show

Benton's statements about possible superintelligence, loss of control, and existential consequences are his assessment. They do not establish that an intelligence explosion has occurred, that loss of control is imminent, or any specific probability. His move to METR is not independent confirmation of his risk forecast. The signal does not affect incident counts or the timeline and can contribute only through the bounded context factor in the Loss of Control Score.

Classification: This page records a sourced warning, research finding, policy action, or industry response. It is deliberately separate from the Incident register and does not prove that an unauthorized AI action occurred.