The AI Agent Copied the Attack Into Its Own Reply
In a simulated email task, an AI assistant copied an attacker-written instruction into its reply. The instruction posed as a filing rule inside the message it had been asked to answer. OpenAI says no impact occurred outside training and evaluation. The trace shows one hand-off, not an outbreak.
What the test established
The attacker controlled the incoming message. The defender completed the scheduling request but also quoted the whole email, preserving the planted text. The disclosed email and file tests used internal GPT-5.4-mini-based checkpoints; a separate test involved GPT-5.5. The 25 September report names no affected deployed product, success rate or real-world spread. OpenAI plans to train against self-reproduction, without claiming a universal fix.
A copied message is a new security boundary
The practical risk is easy to miss because the first agent can appear to finish its assignment. It may send a useful reply, write a plausible note or commit a reasonable change while also carrying hostile instructions into the next place people or systems will read. The dangerous output can be a quotation, a summary, a file comment or a chat message. Ordinary business workflows already move those forms of text between inboxes, document stores and repositories.
For a prompt to propagate, three things must line up. An attacker needs influence over material the first agent reads. That agent needs permission to write to a shared destination. A later agent needs to encounter the copied material and mistake it for an instruction it should follow. The OpenAI email trace establishes the middle hand-off under test conditions. Whether a second agent would obey depends on the receiving workflow and its controls. Sending a copy is therefore a measurable warning sign, not proof of sustained spread.
The model and its surrounding application have different jobs in this failure. The model treated lower-trust source text as a direction. The application made that decision consequential by allowing the agent to send a message. An organisation cannot assess its exposure from the model name alone. It needs to inspect where each agent gets content, which identity it uses to write, and which other automations consume its output.
This is a known research direction, not the first appearance of an AI worm idea. Morris II research published in 2024 evaluated self-replicating prompts across a controlled ecosystem of retrieval-based email assistants, including multi-hop data extraction. That study and OpenAI’s newer internal test use different environments. Together they justify testing propagation paths in deployed workflows without claiming that either paper documents an outbreak across ordinary customer systems.
Test the whole route, not just the first answer
BlackTree recommends a harmless rehearsal for organisations that let agents read external email, web forms, chat, documents or code, then write to a shared location:
- Draw the read-to-write map. For each agent, record untrusted inputs, output destinations, tool permissions and the people or agents that consume those outputs. Include automatic email quotations, summaries, issue comments and generated files.
- Seed an inert marker. Put a distinctive, non-instructional string in a test input. Run the legitimate task and record whether the string is copied into a reply, file or post. Then let a second test agent process that output. A marker measures movement; it does not require a live malicious payload.
- Check the decision point. If the agent can send or modify consequential content, the application should expose the destination, proposed content and source item to an approver. The approval check should be enforced outside the model, with an audit trail tied to the agent’s identity.
- Constrain and watch the write path. Grant only the send and edit rights the task needs. Log unusual fan-out, repeated text across channels and shared artefacts that acquire instruction-like content from external sources.
- Practise containment. Be able to suspend the agent’s write access, locate downstream copies and pause automated readers of suspect material while an investigation establishes which actions were actually taken.
These are BlackTree’s operational deductions, not controls evaluated by the September OpenAI test. OpenAI’s separate prompt-injection guidance likewise recommends limiting agent access and reviewing important actions before confirmation. The rehearsal needs the organisation’s real connectors and permission settings because a model-only prompt test cannot reveal what an email or repository tool will permit.
Procurement teams can turn this into a concrete request. Ask a supplier to show how its agent treats quoted external instructions, how it distinguishes source text from authority, what an approver sees before a send, and whether audit records connect the input to the output. Ask for a test involving a downstream agent, not only a screenshot of the first agent refusing one suspicious sentence. No single refusal result establishes that copied content cannot travel through another route.
BlackTree previously covered a defensive context-bomb experiment, where a planted instruction stopped an attacking agent, and the broader question of how much authority agents should receive. This disclosure adds a focused operational question: can one agent’s output become the next agent’s instruction? For write-capable deployments, the answer should come from a controlled test of the full workflow.
Sources
- OpenAI Alignment, “Self-replicating prompt injections exist”, discovered 27 June 2026, disclosed and updated 25 September 2026. The page gives no publication time.
- Cohen, Bitton and Nassi, “Here Comes The AI Worm”, first posted 5 March 2024.
- OpenAI, “Understanding prompt injections”, guidance checked 28 September 2026. The reviewed page does not expose a publication date.


