OWASP LLM06 Sensitive Information Disclosure: The Output-Side Controls a Gateway Enforces
OWASP LLM06 covers sensitive information disclosure: the model emits data the application or the user is not authorized to receive. The disclosure paths split into three: training-data leakage, in-context leakage from RAG and tool outputs, and cross-tenant leakage from shared deployments. The output-side controls live at the gateway, where every response is observable before it reaches the user. This article walks through the LLM06 disclosure paths, the output-side controls that work, and the redaction and policy patterns to enforce.

OWASP LLM06 covers sensitive information disclosure. The OWASP Top 10 for LLM Applications defines the failure as insufficient protection against sensitive information in model outputs. For the common training question, the correct answer is the option where an LLM reveals confidential data, PII, or proprietary information in its response. The disclosure paths divide into memorized training content, sensitive context brought in by retrieval or tools, and data crossing tenant boundaries through shared deployment components.
The output-side controls live at the gateway. The gateway sees every response before it reaches the user and can classify it, redact matched sensitive elements, block delivery, or apply a stricter response policy based on the calling identity. An application can filter output too, although a bug or bypass in that application leaves every other application to implement its own equivalent control.
I want to walk through the three disclosure paths, the output-side controls that work at the gateway, the redaction and routing patterns that hold under attack, and the residual application work that the gateway cannot replace.
TL;DR
LLM06 means an LLM response exposed confidential data, PII, or proprietary information. A response proxy can inspect and enforce policy on routed model traffic, while retrieval scoping, database authorization, and cache design remain application controls.
The three disclosure paths
The disclosure paths look similar on the wire and require different controls.
Training-data leakage occurs when a model emits content it memorized during training. The content can be PII, proprietary text, or system-level data that entered a training corpus. The model's training process produced an association between a prompt pattern and that content.
In-context leakage occurs when the current session supplies sensitive material through retrieval, tool output, or system-prompt content. A RAG system may retrieve a document containing PII, then the model repeats it in a summary. A tool can return a database row whose sensitive columns end up in the model response. That path originates in the application's data flow.
Cross-tenant leakage occurs when a retrieval layer or cache confuses one tenant's data with another's. A shared RAG index without tenant scoping can return tenant A's documents for tenant B's query. A response cache keyed only on prompt content can replay a tenant A response to tenant B. The deployment architecture created that exposure.
The output-side controls at the gateway
Six output-side controls do most of the LLM06 work when enforced in the response proxy.
Per-identity PII classification on responses. The proxy runs the response through a classifier that flags PII patterns. Identified PII is redacted, the response is blocked, or the response is routed to a stricter policy depending on the calling identity's data scope. The classifier covers names, email addresses, phone numbers, government identifiers, financial account numbers, health identifiers, and any custom patterns defined by the deployment.
Per-identity content-class filtering. Beyond PII, the gateway can flag content classes the calling identity is not authorized to receive: legal advice, medical advice, internal pricing, source code, and records belonging to another customer. A support-agent identity can receive a sanitized response while a customer identity receives only records assigned to that customer.
Schema-bound responses for tool-mediated outputs. When the response is the output of a tool invocation, the gateway can enforce that the response shape matches the declared schema. A tool that is supposed to return a customer record's first name and last login returns exactly those fields; any additional fields the tool emits are stripped at the gateway. The control protects against tools that return more than they advertise.
Per-tenant response policy at the proxy. The proxy can bind the authenticated caller to an application-supplied tenant value and apply output rules for that tenant. Retrieval-index filters, database authorization, and cache-key design still belong in the application. The response-boundary check cannot prove that a retrieved document had the correct tenant scope without trustworthy provenance in the request or response.
Output-length and content-shape caps. The gateway can enforce maximum response lengths, maximum number of returned records, and structural caps on how much of a retrieved document gets echoed back. The caps reduce the surface for bulk-exfiltration patterns where an attacker tries to extract a large corpus through a single response.
Audit on every response with the classifier result. Every response produces a per-decision audit record that includes the response classifier outcome, the redactions applied, the policy decisions, and the calling identity. The audit record is the forensic trail for any post-incident investigation of a disclosure event.
The redaction and routing patterns
Three patterns recur in production deployments.
Redact-and-pass. PII patterns are replaced with placeholders in the response before it reaches the user. The response is logged with both the original and the redacted form for audit. The user receives a response that still answers the question but with the sensitive fields scrubbed. The pattern works when the response value to the user is in the structure, not the specific PII.
Block-and-route. When the classifier hits a high-severity content class, the response is blocked entirely and the user receives a generic refusal. The original response is logged for review. The pattern is used when redaction is unsafe (the response leaks information just by existing) or when the calling identity is not authorized to receive any portion of the content class.
Route-by-identity. The same prompt routes to different policy paths based on the calling identity. A privileged identity may receive the full response, while a restricted identity receives a redacted version. The pattern enables one application to serve multiple authorization classes without separate per-class deployments.
What sits outside the gateway boundary
The model's training history is outside the gateway boundary. The gateway cannot un-train data that was memorized. The gateway can only detect and redact the memorized content as it surfaces in responses.
The application's own data-flow design is outside the gateway boundary. If the application chooses to push sensitive context into the prompt without scoping it to the user's authorization, the gateway can detect and redact, but the architectural fix is in the application. The gateway is the safety net; the application is the primary control.
This is one of the cases where the DeepInspect HTTP-boundary rule applies cleanly. If the leakage path is the application reading data from a database and writing it to a file the user has access to, the file write never touches the gateway. The gateway sees the AI HTTP traffic, not the application's file system or database operations.
Evidence and compliance context
GDPR Articles 5 and 32 require appropriate security for personal data. A model response that exposes PII to an unauthorized recipient creates a disclosure event. Output classification and a preserved policy decision can support the controller's technical-measures evidence. OWASP AISVS 1.0 is a useful verification reference for output validation and data-protection controls.
HIPAA's Privacy Rule restricts disclosures of PHI to the minimum necessary for the disclosure's purpose. A model that surfaces PHI in a response to a user not authorized to receive it is a Privacy Rule violation. The redact-and-pass pattern aligned with the minimum-necessary standard is the architectural answer.
EU AI Act evidence requirements depend on the system role and classification. A gateway decision record can document what crossed the routed HTTP model boundary, its classification result, the applied policy, and the caller identity. It cannot replace the provider's system-wide technical documentation or the application's own records of retrieval and downstream actions. EU AI Act Article 12 logging explains the broader record-keeping boundary.
DeepInspect
This is the output-side control DeepInspect provides for the LLM06 surface. DeepInspect sits inline between authenticated users or agents and the LLMs they call, classifies every request and response against PII and content-class policies, redacts or blocks identified disclosures based on the calling identity's authorization, enforces per-tenant scoping as a second line of defense, and writes a per-decision audit record outside the calling application.
The classifier runs before the response reaches the user. An application bug that would have surfaced PII to an unauthorized recipient can be redacted or blocked at that point. The policy check evaluates the tenant value supplied with the authenticated request, while application controls retain responsibility for retrieval and cache scoping. The audit record produces evidence for post-incident review.
If output filtering relies on each application team rebuilding the same control, the gap will show up in an exception path. Let's talk today.
Frequently asked questions
- Can the gateway catch every PII leak?
No classifier achieves 100% recall. The gateway's classifier catches the documented PII patterns and any custom patterns the deployment defines. False negatives happen and are part of the residual risk the architecture has to plan for. The defense in depth includes upstream controls (do not put PII in prompts that do not need it) and downstream controls (audit log for post-incident review).
- How is LLM06 different from LLM02 (insecure output handling)?
LLM02 covers downstream code-execution risks from the model's output: an output that contains SQL the application executes, JavaScript the browser renders, or shell commands a script invokes. LLM06 covers content the output should not have contained in the first place. The two overlap when the sensitive content is also dangerous to execute (a leaked API key the application then uses).
- What does cross-tenant leakage look like in practice?
Documented cross-tenant cases often involve shared retrieval indexes, response caches, or session stores. A multi-tenant RAG system with an unscoped retrieval filter can return one tenant's documents for another tenant's query. Per-tenant scope belongs at every shared resource; LLM response schema validation covers an adjacent output-side check.
- Does the classifier slow responses down?
Modern PII classifiers run in low-millisecond regimes and the work is dominated by network round-trips, not classification. The end-to-end latency overhead is invisible relative to the model's inference time.
- Can the model itself be instructed not to leak?
Marginally. Instructions in the system prompt help against accidental leakage and do little against adversarial extraction. The instruction is worth including; it cannot be the only control.
- Where does this fit in OWASP AISVS?
OWASP AISVS Chapter 6 (output handling) and Chapter 8 (data protection) cover the verification requirements for the LLM06 surface. The chapters require documented output classification, redaction policies, per-tenant scoping checks, and per-decision logs of classifier outcomes.