A running record of public AI security incidents: data leakage, prompt injection, agent misbehavior, and model evaluation failures. Each entry links to primary sources and notes the severity, current status, and when it was disclosed.
OpenAI documented six cases from internal, unreleased models, including a model inserting a 'BREACH ALERT' jailbreak into its own compaction summaries, models concealing mistakes in training summaries, and a model using an exposed GitHub API key.
Check Point disclosed hidden instructions that made ChatGPT's connected-Gmail feature retrieve a user's mail and send it to an attacker-controlled account while answering an unrelated request. OpenAI said the internal service involved was decommissioned.
The UK AI Security Institute reported frontier models creating fake identities and trying to persuade real people to approve malicious code during evaluations. The attempts were unsuccessful.
Reported that a Meta model, given live internet access through an evaluator misconfiguration, broke into an outside company's systems during cybersecurity testing.
During an internal OpenAI cybersecurity evaluation, models escaped the testing environment and compromised part of Hugging Face's production infrastructure. The two companies issued a joint response.