Former Anthropic safety researcher Joe Benton warns about loss of control and calls for transparency
Joe Benton said he left Anthropic because, in his view, leading AI companies underinvest in safety and do not give the public enough visibility into loss-of-control risks, safety incidents, or progress toward self-improving systems. He is joining METR to support independent evaluations. His statements are an expert warning, not empirical evidence that loss of control will occur.
- Published
- Sep 11, 2026
- Signal type
- Insider warning
- Confidence
- 99%
- Organization
- Anthropic, METR
- Sources
- 3
- Last reviewed
- Sep 14, 2026
Assessment
What this signal changes
Benton says he led a scalable-oversight team at Anthropic and describes a structural conflict between the race for more capable systems and adequate safety investment. He calls for reporting of incidents and near misses, minimum standards, and independent assessments. His original essay, personal website, and independent Associated Press reporting establish his role, move to METR, and core statements.
Evidence boundary
What the evidence does not show
Benton's statements about possible superintelligence, loss of control, and existential consequences are his assessment. They do not establish that an intelligence explosion has occurred, that loss of control is imminent, or any specific probability. His move to METR is not independent confirmation of his risk forecast. The signal does not affect incident counts or the timeline and can contribute only through the bounded context factor in the Loss of Control Score.