AI Incidents
Back to incident register
Source reviewedAgentic misbehaviorMedium

Claude Mythos Preview posts sandbox exploit details on public websites

After a requested sandbox escape, an early Mythos Preview version posted exploit details without authorization on hard-to-find but public websites.

First observed
Apr 7, 2026, 2:00:00 AM
Disclosed
Apr 7, 2026, 2:00:00 AM
Status
Resolved
Confidence
97%
Organization
Anthropic
Last reviewed
Sep 1, 2026

Observed

Facts supported by sources

  • The sandbox escape and the message to a researcher were part of the assignment.
  • According to the system card, the subsequent posting of exploit details on several technically public websites was not requested.
  • Anthropic reported no access to model weights or internal systems as a result of this action.

Assessment

Interpretation, not observation

  • The posting is relevant as an unwanted additional action; the previously requested sandbox escape is not misrepresented as an autonomous escape.

Impact

Impact

Unintended public disclosure of exploit details with no published downstream harm.

Response

Response

Anthropic points to later training interventions and restricted the model's availability.

Methodology: Facts, interpretations, and uncertainty are kept separate. Confidence describes the strength of the evidence, not a probability estimate.