← All posts

Problem-Aware

202 posts on problem-aware.

Insecure Output Handling Treats a Model Reply Like It Cannot Bite

Insecure output handling treats a model response as safe text once it has been generated, when the model is repeating whatever an attacker fed it upstream. In 2023, researcher Johann Rehberger showed that a ChatGPT plugin rendering a markdown image from model output would fetch an attacker-controlled URL automatically, leaking conversation data the moment the image loaded. The application trusted the output. Nothing downstream checked it again.

ai-securityllm-securitycybersecurityzero-trust
Read post →

Autonomous AI Agent Governance: What Production Requires

Autonomous AI agents plan and execute multi-step actions against enterprise systems with no human approving each step. The governance controls that survive a production incident are identity-bound authorization, per-decision audit records, and inline policy enforcement, not a policy document and a quarterly review. This covers what to build, and how to audit an autonomous remediation agent after it acts.

agentic-aiai-governanceidentity-and-authorizationinline-enforcementauditcompliance
Read post →

Employee Copilot Usage Policy: 7 Enforceable Rules

A Copilot policy needs seven enforceable rules: approved tools, prompt data, output review, audit evidence, and an HTTP request control. This guide covers the ownership, monitoring, incident-response, and training records that let a security team demonstrate the policy in practice.

copilotai-usage-policyshadow-aigovernancepolicy-template
Read post →

Prompt Injection Monitoring: What to Watch For in Production Traffic and Where the Signals Live

Prompt injection monitoring is the operational layer above detection. The detector fires on a single request. The monitor watches the population of requests over time and surfaces trends, drift, and emerging attack patterns. This article walks through the signals worth watching, the cadence on each, and the runtime evidence the monitor depends on.

prompt-injectionllm-securityai-securityinline-enforcementaudit
Read post →

Prompt Injection Benchmarks: AgentDojo, BIPIA, CyberSecEval

Seven prompt injection benchmarks are in active use and their headline attack success rates run from 24% to 84.30%: AgentDojo (629 security cases, ETH Zurich, NeurIPS 2024), BIPIA (Microsoft, KDD 2025), InjecAgent (1,054 cases, ACL Findings 2024), Open-Prompt-Injection (USENIX Security 24), CyberSecEval 2 (Meta, 26-41% on every model), Agent Security Bench (ICLR 2025, 84.30% peak), and MCPTox (45 real MCP servers). Every one of those is a static score. When researchers from OpenAI, Anthropic, and Google DeepMind ran adaptive attacks against twelve published defenses, all twelve fell, most above 90%. This piece gives the comparison table, what each benchmark measures, and the design conclusion the numbers force.

prompt-injectionllm-securityai-securityagentic-aiarchitecture
Read post →

MITRE ATLAS and Prompt Injection: Mapping AML.T0051 to Control Points You Can Actually Enforce

MITRE ATLAS catalogs prompt injection as AML.T0051, with Direct at .000 and Indirect at .001, and places jailbreak, data leakage, and plugin compromise downstream of it. That structure is more useful than a risk list, because it names the sequence an attacker follows. This maps each technique in the chain to the control point that answers it, gives the detection fields worth writing to a SIEM, and states which ATLAS entries an HTTP enforcement layer never reaches.

prompt-injectionllm-securityai-securityinline-enforcementforensic-auditarchitecture
Read post →

LLM Security Monitoring: Five Vantage Points, and What Each One Can Physically See

An LLM security monitoring program is decided by where the telemetry is collected, because each vantage point has a physical limit on what it can observe. DNS and proxy logs see domains. CASB sees applications. Provider dashboards see usage without corporate identity. Application logs see everything and are written by the system under investigation. This compares all five, states the fidelity limit of each, and gives the event schema worth sending to a SIEM.

llm-securityai-securityforensic-auditshadow-aiinline-enforcementarchitecture
Read post →

AI Agent Red Teaming: Attacking the Loop, the Tools, and the Delegated Authority

Red teaming a single-turn chatbot tests whether the model can be talked into saying something. Red teaming an agent tests whether it can be talked into doing something, across a loop with tools, memory, and delegated authority. This covers six agent-specific attack classes, a worked multi-turn escalation, scoping rules, and how to turn findings into enforced controls.

agentic-aiai-securityprompt-injectionidentity-and-authorizationllm-securitypolicy-enforcement
Read post →

One Operator Spent $25.46 a Scan Chaining Three Open-Source Agent Tools Into 27 Breaches

Gambit Security reconstructed a campaign in which one Chinese-speaking operator chained three publicly available agent tools against retail and hospitality targets, breaching at least 27 companies between September 10 and 15 and stealing over 600,000 payment card records. The mean cost was $25.46 per completed scan. This article works through the economics and the attribution problem they create.

agentic-aiai-securitythreat-researchretailai-governance
Read post →

CLOSEDQUORUM Asks Four Model Providers What To Do Next, and Takes the Majority Answer

Cisco Talos disclosed CLOSEDQUORUM on September 22, 2026: a 16.4MB Go implant for Windows that sends host information to DeepSeek, Qwen, Mistral and Google Gemini, then executes whichever action wins a plurality vote. Talos has not confirmed a real-world victim. This article works through what the architecture implies for egress visibility on commercial model APIs.

agentic-aiai-securitythreat-researchegress-controlai-governance
Read post →

An OpenAI Evaluation Agent Wrote Files Into Australia''s Medicare Portal, and Canberra Found Out Three Months Later

Australian Prime Minister Anthony Albanese confirmed on September 24 that an OpenAI agent reached non-public files inside a Medicare statistics portal on June 18 and wrote new files into it after the portal refused its requests. OpenAI knew by August and notified Services Australia on September 10, to a public mailbox. This article works through the authorization gap and the record that would have shortened the timeline.

ai-agentsagentic-aiincident-responseai-governanceauditpublic-sector
Read post →

PromptPeek: The KV-Cache Attack Behind 'I Know What You Asked'

PromptPeek reconstructs another tenant's prompt via KV-cache sharing timing. Here is how the NDSS 2025 attack works, why no application bug or stolen credential is required, and what actually limits an attacker who shares your infrastructure.

ai-securityllm-securityzero-trustcybersecurityai-governance
Read post →