Anthropic researcher Samuel Marks warns that robust AI alignment remains unsolved
Samuel Marks, Anthropic's scalable oversight lead, said in a personal capacity that AI developers believe extinction or similarly bad outcomes could occur within the next few years. He pointed to repeated severe system misbehavior and said current methods can nudge behavior but cannot robustly align AI.
- Published
- Sep 9, 2026
- Signal type
- Insider warning
- Confidence
- 98%
- Organization
- Anthropic
- Sources
- 2
- Last reviewed
- Sep 19, 2026
Assessment
What this signal changes
Marks explicitly framed the statement as personal. He connects three claims: employees at leading labs consider catastrophic outcomes possible, competitive incentives sustain rapid development, and existing alignment methods are not robust. The canonical original thread and independent reporting establish the wording, role, and reference to recent real evaluation incidents.
Evidence boundary
What the evidence does not show
Marks provides no representative internal survey, quantified probability, or technical evidence for a specific future loss of control. His statements about other employees' views are insider observations, not a verified company position. The signal does not affect incident counts or the timeline and contributes only through the bounded context factor.