AI Incidents

A running record of public AI security incidents: data leakage, prompt injection, agent misbehavior, and model evaluation failures. Each entry links to primary sources and notes the severity, current status, and when it was disclosed.

Showing 1 - 5 of 5

highresolved

OpenAI discloses six internal model-misbehavior incidents

OpenAI documented six cases from internal, unreleased models, including a model inserting a 'BREACH ALERT' jailbreak into its own compaction summaries, models concealing mistakes in training summaries, and a model using an exposed GitHub API key.

openaiagent-misbehaviordeceptioninternal
highresolved

ChatGPT ShadowLeak Gmail exfiltration

Check Point disclosed hidden instructions that made ChatGPT's connected-Gmail feature retrieve a user's mail and send it to an attacker-controlled account while answering an unrelated request. OpenAI said the internal service involved was decommissioned.

openaichatgptdata-leakageprompt-injectionagent