AI Incidents
Back to published register
Risk signalInsider warningNot an incident

Anthropic researcher Samuel Marks warns that robust AI alignment remains unsolved

Samuel Marks, Anthropic's scalable oversight lead, said in a personal capacity that AI developers believe extinction or similarly bad outcomes could occur within the next few years. He pointed to repeated severe system misbehavior and said current methods can nudge behavior but cannot robustly align AI.

Published
Sep 9, 2026
Signal type
Insider warning
Confidence
98%
Organization
Anthropic
Sources
2
Last reviewed
Sep 19, 2026

Assessment

What this signal changes

Marks explicitly framed the statement as personal. He connects three claims: employees at leading labs consider catastrophic outcomes possible, competitive incentives sustain rapid development, and existing alignment methods are not robust. The canonical original thread and independent reporting establish the wording, role, and reference to recent real evaluation incidents.

Evidence boundary

What the evidence does not show

Marks provides no representative internal survey, quantified probability, or technical evidence for a specific future loss of control. His statements about other employees' views are insider observations, not a verified company position. The signal does not affect incident counts or the timeline and contributes only through the bounded context factor.

Classification: This page records a sourced warning, research finding, policy action, or industry response. It is deliberately separate from the Incident register and does not prove that an unauthorized AI action occurred.