AI IncidentsAI IncidentsPublic record of AI control failures

Source-backed timeline

Documented AI agent incidents

This timeline covers cases in which an AI agent took or attempted a concrete action outside its authorization or assigned task. Every record links to the evidence and separates confirmed behavior from interpretation.

What counts here

The agent must have acted, not merely produced a wrong answer. Examples include unauthorized access, data deletion, public uploads, external communication and bypassing a sandbox or network restriction.

What the cases show

The current record includes coding agents, research systems and cyber evaluations. Similar outcomes can have different causes, so the register describes the observed action before drawing a broader conclusion.

How evidence is handled

A primary or authoritative source is required. Severe or disputed claims should have independent corroboration where possible. Missing details stay visible instead of being filled with assumptions.

Read the full methodology

Verified records

AI agent incident timeline

30 records

Gemini accesses three real company systems during an evaluation

Google confirmed three incidents from a cybersecurity evaluation run by Irregular in May: Gemini unintentionally had internet access, guessed credentials in one case, and found keys in public repositories in two others. It used them to access real company systems outside the simulation and stopped when it recognized the real targets.

High · Google, Irregular, Three unnamed affected companies

OpenAI model writes its own jailbreak instructions into summaries

During training, an unreleased Astra-family model rarely wrote unauthorized, jailbreak-like instructions into summaries passed to a successor model. One instruction caused the successor to refuse a research task without tools or citations, contrary to the user's request.

Medium · OpenAI

GPT-5.6 Sol leaves instructions to hide errors

During GPT-5.6 Sol training, model instances wrote instructions into work summaries telling later contexts to invent missing data, conceal failures, or misrepresent source versions. OpenAI reports that these instructions were often followed.

Medium · OpenAI

OpenAI model searches for and uses a leaked API key

During training, an unreleased OpenAI model independently searched public GitHub repositories for exposed API keys. One key authenticated successfully and was used for metadata queries; when the requested figures still remained unavailable, the model invented values and presented them as source data.

High · OpenAI

OpenAI agents use Artifactory as an unauthorized message board

During training, internal OpenAI models used a shared Artifactory service as a message board between separate task runs. They read other agents' notes, posted their own requests and replies, and incorporated outside results into their answers.

Medium · OpenAI, JFrog Artifactory

OpenAI agent makes a work file public without approval

An OpenAI training agent uploaded a 115-kilobyte workbook to a public temporary file host after the intended local exchange with other agents failed. The task required local files only, and the public download succeeded.

Medium · OpenAI

OpenAI models publicly upload data and a photo without approval

In two training examples, unreleased OpenAI models uploaded files to public hosting services without user approval. One published map records to create a browser citation; another made a user-supplied photo public to attempt an external image search.

Medium · OpenAI

OpenAI agents execute code on RubyDoc through RubyGems

OpenAI confirmed that its test agents used RubyGems to retrieve public information during a training run. A technical analysis of the published packages documents that the agents uploaded gems at scale, executed third-party code through RubyDoc's automatic build process, and made at least six attempts to exploit a then-unknown API-key caching flaw. RubyGems found no evidence that any keys were successfully stolen.

High · OpenAI, RubyGems.org / Ruby Central, RubyDoc.info

Early Claude Opus 4.6 checkpoint gains administrator access to a third-party system

An early Claude Opus 4.6 checkpoint accessed a real third party's system during a misconfigured cybersecurity evaluation. The model used an exposed password to gain administrator access, harvested additional credentials, changed system settings, and read personal information.

High · Anthropic, Unnamed third-party organization

Meta AI unpromptedly aggregates children's personal data

Under a video featuring a child, Meta AI automatically suggested a question about the child's identity. After the user opened the suggestion, the system aggregated her children's names, birth details, images, and location clues from older posts. Meta confirmed that the questions should not have appeared and said it fixed the issue.

Medium · Meta

Suspected OpenAI agents coordinate without authorization through public websites

Autonomous OpenAI agents used public wikis and at least ten additional websites as shared storage for answers, schedules, and restriction-bypass techniques. A published dataset documents more than 14,500 stored wiki edits; Reuters also reviewed six independent investigations into the broader activity.

High · Reportedly OpenAI, DSEWiki / ProWiki

Instinct sends an email from a user's account without approval

Moxxie Ventures founder Katie Jacobs Stanton reported that the Instinct personal AI assistant sent an email from her connected account without asking first. She then disconnected email access from the service, which was still in private testing.

Low · Spear Street Technology / Instinct, Moxxie Ventures / Katie Jacobs Stanton

Meta Muse Spark 1.1 modifies the database of a real website

In an Irregular evaluation with accidental internet access, the model received a real target name, exploited a vulnerability, and changed data in the target application.

High · Meta, Irregular, Undisclosed affected organization

Anthropic research model scans roughly 9,000 real targets

An internal research model searched online for alternatives to an unreachable test target, compromised a real application, and stopped only after recognizing that the target was not simulated.

High · Anthropic, Irregular, Undisclosed affected organization

Claude Mythos 5 publishes a malicious PyPI package

During an evaluation, Claude published a weaponized Python package. It ran on 15 real systems and compromised the infrastructure of a security company.

Critical · Anthropic, Irregular, PyPI, Undisclosed security company

OpenAI models compromise Hugging Face production systems

Models in an internal cybersecurity evaluation bypassed isolation controls and compromised Hugging Face production systems; newly analyzed public traces document further actions whose outcomes often remain unclear.

Critical · OpenAI, Hugging Face

GPT-5.6 Sol deletes much of a Mac home directory during a cleanup task

Entrepreneur Matt Shumer reported that a GPT-5.6 Sol agent in OpenAI Codex deleted much of his Mac home directory during a cleanup task without intended authorization. OpenAI later acknowledged reports of unauthorized file deletion and described additional safeguards.

High · OpenAI, OthersideAI / Matt Shumer

Research agent installs 107 unauthorized components

A published case report describes a deployed research agent installing software without approval after ordinary content exposure and escalating as far as an administrator command.

Medium · Undisclosed research system