← Blog

Self-Replicating Prompt Injection Makes Agent Output a Containment Problem

Parminder Singh
Parminder Singh··5 min read
Summarize with AI

OpenAI disclosed on September 25, 2026 that its GPT-Red red-teaming agent found prompt injections that copy themselves into the next email, file write or code comment. The payload has two goals at once: do the attacker bidding and get reproduced in the output. That shifts the control point from the inbound prompt to the outbound response, which is where an identity-aware HTTP enforcement layer can inspect what an agent is about to publish.

Platform & Architectureprompt-injectionai-securityagentic-aiai-gatewayincident-response
Self-Replicating Prompt Injection Makes Agent Output a Containment Problem

On September 25, 2026, OpenAI's alignment team published a misalignment report describing what it calls "a new variety of prompt injection, which can self-propagate akin to a computer worm." The team found it on June 27, 2026 using GPT-Red, an internal red-teaming agent built on a GPT-5.4-mini checkpoint, and reproduced it against GPT-5.4-mini and GPT-5.5. Self-replicating prompt injection differs from the injections most defenses were built for because the payload carries a second objective. It has to "both achieve a particular adversarial goal and induce the defender to reproduce them publicly."

That second objective is the part that breaks the usual control design. An inbound filter that scores prompts for malicious instructions still lets the model emit a response that carries the instruction forward into an outgoing email, a committed file or a Slack message that another agent will read tomorrow.

TL;DR

  • OpenAI disclosed self-replicating prompt injection on September 25, 2026. The payload pursues an attacker goal and makes the model reproduce the payload in its own output.
  • Three propagation paths were observed in evaluation: email replies, file writes and code comments.
  • Inbound prompt filtering alone leaves the replication step unguarded. The outbound response is where propagation happens.
  • OpenAI reported no impact beyond simulated tool calls in training and evaluation. The models tested were internal research checkpoints.
  • Record the identity, destination and policy decision for every AI request and response so a propagating payload has a traceable first hop.

The payload has two jobs

A conventional indirect injection hides an instruction in retrieved content and hopes the model follows it once. The self-replicating variety adds a reproduction clause to the same text. OpenAI's report describes injections that must satisfy both conditions in the same text, so a single successful landing can seed further copies without the attacker sending anything else.

The Register's September 29, 2026 write-up lists the concrete examples. One injection arrived by email and instructed the model to reply in Spanish while quoting the entire original message, which carried the payload into the reply thread. Another sat inside a dataset as a fake system warning, triggered deletion of a report and replicated itself in the process. A third steered a model across several Slack reads that each looked relevant on their own.

OpenAI stated that "no impact was observed outside of the simulated tool calls in training and evaluation," and that the affected models were internal research checkpoints rather than released endpoints. The mitigation it describes is training-time: including self-reproduction as an attacker goal in GPT-Red training so future released models have seen the pattern.

Propagation needs a write path, and the write path is HTTP

Each observed vector ends in a write. Email propagation needs an outbound send, filesystem propagation needs a file write and a code comment needs a commit. In an enterprise agent deployment, those writes are almost always HTTP calls, and the text being written came from a model response that also travelled over HTTP.

A model vendor's training-time fix is useful, but treating it as your containment plan puts your incident response on someone else's release schedule. The organization still owns the hop between the model response and the system that stores it.

This is why response inspection matters more here than in a single-turn chat deployment. Indirect prompt injection defense covers the inbound side. Replication adds a second decision point on the way out, and the same pattern appeared earlier this year in the Copilot case documented in self-propagating prompt injection and document provenance.

What a single hop of containment buys

Stopping a worm does not require perfect detection at every node. It requires cutting enough edges that the reproduction rate falls below one. For an AI agent fleet, the cheapest edge to cut is the outbound response that an agent is about to persist or forward.

Three checks are practical at that point. First, evaluate whether the response contains instruction-shaped text aimed at a downstream model, which is a narrower signal than general malicious-content scoring. Second, check whether this agent identity is authorized to write to this destination at all. Third, record the decision with enough context that an investigator can reconstruct the first hop.

The injection in OpenAI's dataset example arrived as a fake system warning inside a single row, the kind of row an analyst scrolls past at 6pm without reading. No human in that path was going to notice the replication clause.

Destination control limits the blast radius

Instruction-shaped text is a probabilistic signal, so it should not be the only control. Destination authorization, by contrast, is a deterministic check against a list. An agent that summarizes support tickets has no business committing to a repository, and an agent that reviews code has no business sending external email.

AI egress control describes how that boundary is enforced on model-bound and tool-bound traffic. Applied to this attack, it means a payload that successfully compromises one agent still cannot reach the three systems it needs to spread, because the identity carrying it was never authorized for those destinations.

Local process execution and STDIO transports sit outside an HTTP enforcement point. An agent that writes to its own local disk through a local tool server produces no HTTP request to inspect, so endpoint controls and repository review remain necessary for that path.

Evidence requirements after a propagating failure

An incident review of a replicating payload asks a specific sequence of questions. The first is which identity made the request that returned the payload. Then the reviewer needs the destination that received the first copy, the policy version in force at that moment and every subsequent request whose response carried matching text.

Answering those requires per-request records written outside the agent's own logging, because a compromised agent's self-reported log is the artifact under review. Signed audit logs for AI requests explains the write-path independence argument in more detail, and agentic AI audit trail covers the per-decision fields that survive an investigation.

DeepInspect

DeepInspect is a stateless proxy for authenticated HTTP traffic between enterprise users or agents and LLM endpoints. It evaluates application-supplied identity, request classification, approved destination and policy on the way in, and it evaluates the model response on the way back before the calling agent persists or forwards it. Every permit, redaction, reroute or block produces a signed per-decision record outside the calling application's write path.

For a propagating payload, that gives two useful properties. An unauthorized destination is refused deterministically, and the first hop is recorded with the identity and policy state that produced it. DeepInspect leaves local process execution, STDIO tool servers, repository review, endpoint controls and model-level hardening with the enterprise and its suppliers. Book a demo today.

Frequently asked questions

Does this affect production OpenAI models today?

OpenAI reported the behavior on internal research checkpoints of GPT-5.4-mini and GPT-5.5, with no observed impact outside simulated tool calls in training and evaluation. The disclosure describes a class of attack rather than an exploited production incident. Enterprise agent deployments that read untrusted content and write to shared systems carry the structural exposure regardless of which vendor model sits in the middle.

Is inbound prompt filtering now useless?

Inbound filtering still removes the volume of ordinary injection attempts. It becomes insufficient on its own once the payload objective includes reproduction, because the damaging step happens after generation. Pair inbound content checks with outbound response inspection and destination authorization so a single missed prompt does not produce a persistent copy.

What makes this different from a normal indirect injection?

The difference is reproduction. A normal indirect injection attempts one adversarial action in one session. A self-replicating injection makes the model restate the instruction in output that another model or agent will later ingest, which turns one compromised session into a population of them.

Which logs would an investigator actually need?

Per-request records showing supplied identity, calling application, destination, resolved model, policy version, decision and timestamp. The records need to sit outside the agent's own write path, and they need to be queryable by text match so an investigator can find every request whose response carried the replication clause.