AI Incidents
Back to published register
Risk signalIndustry responseNot an incident

Anthropic CEO calls for slower AI development and embedded external evaluators

Anthropic CEO Dario Amodei called for slowing the development of the most capable AI models so safety research and control can keep pace. He also announced that Anthropic would give external evaluators employee-like access. His warnings about future loss of control are documented as attributed risk assessments, not as forecasts or a new incident.

Published
Sep 12, 2026
Signal type
Industry response
Confidence
99%
Organization
Anthropic, METR
Sources
3
Last reviewed
Sep 14, 2026

Assessment

What this signal changes

Amodei's original post combines a concrete industry proposal with a verifiable corporate measure: Anthropic says it will give an external evaluation team broad, employee-like access and allow it to publish important safety findings independently. AP and Axios confirm the publication, the call to slow development, and its connection to recent agent incidents.

Evidence boundary

What the evidence does not show

The claims that a similarly misaligned agent swarm could take over the internet within six to twelve months or cause exceptionally large losses are Amodei's personal risk assessment. They are not a measured probability. The original post states only September 2026; AP and Axios establish the specific date as September 12. The signal does not affect incident counts or the timeline and can contribute only through the bounded context factor in the Loss of Control Score.

Classification: This page records a sourced warning, research finding, policy action, or industry response. It is deliberately separate from the Incident register and does not prove that an unauthorized AI action occurred.