OpenAI has generated an internal form of prompt injection that can propagate itself. The attack is designed to do more than trigger an unauthorized action. It also causes the affected agent to republish the malicious instructions in an email, file, code comment or another channel that a later agent may read.
The attacks emerged through self-play between a GPT-Red attacker and a vulnerable model. Both were internal research checkpoints based on GPT-5.4-mini. OpenAI demonstrated simple email examples, filesystem and code variants, and multi-hop attacks in which several messages that appeared ordinary on their own combined to trigger an unauthorized action and retransmit the payload.
One example caused an agent to delete important reports and copy the full attack into a file. In another scenario, messages spread across several Slack channels led an agent to carry out an internal recognition action and repost manipulated text. These are concrete attack demonstrations, not evidence that a customer or a publicly deployed OpenAI product was compromised.
OpenAI explicitly classifies the disclosure as a research finding rather than an incident. The company says every observed effect was confined to simulated tool calls in training and evaluation. The report therefore does not establish a real AI worm or show how often such attacks would succeed outside a controlled environment.
OpenAI is adding self-reproduction as a distinct attacker goal in GPT-Red training so future released models will have encountered these patterns during training. Independent evidence does not yet show how reliably that approach protects open, heterogeneous agent ecosystems. The finding matters because connected agents can read, change and forward the same content that informs their security decisions.