Insecure Output Handling Treats a Model Reply Like It Cannot Bite
Insecure output handling treats a model response as safe text once it has been generated, when the model is repeating whatever an attacker fed it upstream. In 2023, researcher Johann Rehberger showed that a ChatGPT plugin rendering a markdown image from model output would fetch an attacker-controlled URL automatically, leaking conversation data the moment the image loaded. The application trusted the output. Nothing downstream checked it again.