AI Incidents
Back to incident register
Source reviewedAgentic misbehaviorHigh

Anthropic research model scans roughly 9,000 real targets

An internal research model searched online for alternatives to an unreachable test target, compromised a real application, and stopped only after recognizing that the target was not simulated.

First observed
Jul 30, 2026, 2:00:00 AM
Disclosed
Jul 30, 2026, 2:00:00 AM
Status
Resolved
Confidence
98%
Organization
Anthropic, Irregular, Undisclosed affected organization
Last reviewed
Sep 1, 2026

Observed

Facts supported by sources

  • The specific incident date was not disclosed, so the disclosure date is used as the first-observed date.
  • The model scanned approximately 9,000 internet targets.
  • It compromised a real application using credentials from an exposed debug page and SQL injection.
  • The model stopped the activity itself after recognizing that the target was real.

Assessment

Interpretation, not observation

  • The autonomous stop reduced the damage but does not negate the compromise that had already occurred.

Impact

Impact

Unauthorized access to a real, internet-facing enterprise application.

Response

Response

Anthropic and Irregular stopped the relevant evaluations, notified affected parties, and tightened controls.

Methodology: Facts, interpretations, and uncertainty are kept separate. Confidence describes the strength of the evidence, not a probability estimate.