← Blog

LLM JSON Schema Validation: Structure, Policy, and Scope

Parminder Singh
Parminder Singh··6 min read
Summarize with AI

LLM JSON schema validation confirms structure, required fields, and types. This guide shows the separate policy checks for business rules, personal data, authorization scope, and tool-call arguments on an LLM response.

Platform & Architecturellm-engineeringjson-schemaai-engineeringstructured-outputai-security
LLM Response Schema Validation: When JSON Mode Is Not Enough

LLM JSON schema validation through OpenAI Structured Outputs, Anthropic tool use, or Google's response schema confirms a document can parse and meets declared structural constraints. Required fields, field types, and enum values have a defined shape. That is valuable engineering work, especially when an application needs a predictable object rather than prose.

Business policy needs a separate decision. A schema-valid response may contain a discount above an approval threshold, personal data for an unauthorized recipient, or tool-call arguments outside the caller's entitlement. Schema validation verifies the object. Semantic validation evaluates the object against the enterprise rules that apply to this request.

I want to walk through what JSON mode actually covers, the semantic-validation gap it leaves, and the inspection-layer architecture that runs schema validation and semantic validation on the same response path.

TL;DR

  • LLM JSON schema validation checks that a response has the declared structure, fields, types, and constraints.
  • Business rules, data classification, authorization scope, and cross-field checks need semantic validation after parsing.
  • A routed HTTP inspection layer can evaluate both passes before it returns a response to the authenticated caller.

What JSON mode covers

The schema-conditioned response modes across providers deliver a common set of guarantees.

The response parses as valid JSON. There is no trailing prose, no markdown fencing, no malformed structure. Applications that consume the response can call JSON.parse without wrapping the call in defensive text extraction.

The response satisfies the JSON Schema declaration the caller sent: required fields are present, values match declared types, enum values stay inside the allowed list, and nested objects satisfy their nested schemas.

The response respects field-level constraints the schema declares. Minimum and maximum values on numbers. Minimum and maximum length on strings. Pattern constraints on string values via regex.

The response format is deterministic across retries at temperature 0. Two calls with the same prompt and schema produce responses that parse identically (though the content may differ if temperature > 0).

The semantic-validation gap

JSON Schema checks leave four business and security checks for the application or an inline inspection layer.

Business-rule validation

A schema declares that discount_percent is an integer between 0 and 100. The business rule is that discounts above 25 require manager approval. The schema-conditioned response can return discount_percent: 50 and satisfy the schema. The business rule fails, and the application processes the discount unless something else catches it.

Data-classification validation

A schema declares that customer_response is a string. The data-classification policy says customer responses cannot include personal information for accounts in the EU without additional consent. The schema-conditioned response can return a customer response that contains the customer's home address. The schema is satisfied. The classification policy is violated, and the response ships to a channel that has not passed the consent test.

Authorization-scope validation

A schema declares that target_account_id is a UUID. The authorization scope for the current caller is limited to accounts the caller is assigned to. The schema-conditioned response can return any valid UUID, including UUIDs of accounts the caller has no authorization to read or modify. The schema is satisfied. The authorization scope is violated.

Cross-field consistency validation

A schema declares two fields independently. The business rule is that the two fields have to satisfy a joint constraint. order_total = sum(line_items) + tax - discount. The schema declares each field. The schema does not express the arithmetic relationship. The schema-conditioned response can return a document where the arithmetic fails.

Inspection-layer architecture

The four categories of validation run at the inspection layer on the response, before the response returns to the caller. The layer runs two passes over the response.

Pass 1: schema validation

The inspection layer runs the JSON Schema validation the caller declared. The pass catches responses that do not satisfy the declared schema even when the provider's structured-output mode is enabled. Providers occasionally miss the schema constraint (mid-stream errors, provider-side format drift), and the inspection layer's pass is the safety check.

Pass 2: semantic validation

The inspection layer evaluates enterprise semantic rules over the parsed response. Business logic, data classification, authorization scope, and cross-field consistency each attach to the same request context, then apply to every routed response regardless of provider.

A practical policy set can include discount-approval for order.create: when response.discount_percent exceeds 25, it requests manager approval and blocks delivery until that approval exists.

For customer_service.respond, eu-account-response-classification can redact a response and write a classification-violation record when an EU caller lacks a PII-sharing consent flag and the response matches the organization's PII detector.

account-scope-check applies to crm.opportunity.update. When response.target_account_id falls outside caller.assigned_accounts, it blocks delivery and logs the scope violation.

After Pass 1 succeeds, a blocking outcome returns an error without delivering the response. A redaction outcome changes the response before delivery. Approval outcomes route the response to the designated review queue.

The tool-call variant

The same architecture applies to tool calls the model requests. A model that returns a tool_calls array is asking the caller to execute the listed tools. The tool-call arguments are the response equivalent of the JSON body.

The inspection layer runs Pass 1 (schema validation of the tool-call arguments against the tool's declared schema) and Pass 2 (semantic validation of the tool-call arguments against the caller's authorization scope, the target resource's data classification, and the business rules that apply). A tool call that satisfies the schema but violates the scope blocks at Pass 2 before the tool executes.

Compliance implications

The semantic validation layer produces the artifacts multiple compliance regimes require.

The EU AI Act Article 12 record-keeping mandate requires the log to reconstruct what the AI system did. A response that was validated and modified by the inspection layer carries both the pre-validation and post-validation form in the audit record. The auditor can reconstruct the original response, the validation rules that fired, and the modification the inspection layer applied.

The OWASP AISVS 1.0 requirements V3.2 (output validation) and V3.3 (output filtering) map directly to the two-pass architecture. The Pass 1 validation satisfies V3.2. The Pass 2 semantic validation satisfies V3.3.

The GDPR data-minimization principle requires that responses containing personal data are constrained to the identities authorized for the data. The Pass 2 data-classification validation is the enforcement mechanism.

Performance profile

The two-pass validation adds latency to every response. The added latency is dominated by the Pass 2 semantic evaluation, which depends on the complexity of the rules. From internal DeepInspect testing, the added latency is under 10 ms for a typical rule set (10 to 30 rules) on a response under 4 KB. The LLM inference latency itself is 500 ms to 5 seconds, so the added validation latency is under 2% of the total end-to-end response time in the common case.

DeepInspect

This is exactly what DeepInspect does. DeepInspect sits inline between your users or agents and the LLM APIs they call. Every routed response runs through Pass 1 schema validation and Pass 2 semantic validation before it returns to the caller. AI gateway architecture explains where that HTTP proxy sits in the request path.

Semantic rules are declared once at the inspection layer and apply to every routed response. Multiple providers, models, and applications can pass through the same rule set. The audit record captures the pre-validation response, the rules that fired, and the post-validation response the caller received. Signed AI audit logs cover the evidence an investigator needs when a policy decision is challenged.

Book a demo today.

Frequently asked questions

Does the inspection layer break streaming responses?

Streaming responses accumulate into a full response object before Pass 2 evaluates. The client receives the streaming tokens with a small buffer that lets Pass 2 evaluate the complete response before the final chunk closes. For applications where the token-by-token stream is the product experience, the buffer is a tradeoff, and the inspection layer supports policies that allow token-level streaming with a post-stream validation pass for audit purposes only.

How does the layer handle responses that fail Pass 1?

A response that fails Pass 1 is returned to the caller as an error. The audit record captures the model's raw response, the schema that was declared, and the specific validation failure. Applications can implement a retry loop that regenerates with the model and reruns Pass 1. Deployments with strict SLAs cap the retry count and surface the failure to the caller after the cap.

Are the semantic rules declarative or code?

Both patterns are supported. Declarative rules cover the common cases (data classification, authorization scope, simple business logic) and read like the YAML samples above. Complex cases that require custom evaluation logic run as sandboxed code with a defined input contract (the parsed response, the caller identity, the request context) and a defined output contract (allow, block, redact, or approve).

How does semantic validation interact with model fine-tuning?

Fine-tuning changes what the model tends to produce, not what the enterprise's policy allows. A fine-tuned model that is unlikely to violate a semantic rule still runs through the Pass 2 evaluation, because the enforcement layer is where the compliance artifact is produced regardless of the model's training. The inspection layer is the point of record.

Does the layer work with providers that do not offer structured output?

Yes. Providers that do not offer schema-conditioned responses return responses that may or may not parse as valid JSON. Pass 1 validates the parsed content against the schema the caller declared, and responses that fail Pass 1 error to the caller. The inspection layer treats the provider's structured-output feature as a hint, not a guarantee, and validates independently.

What does the audit record show for a redacted response?

The audit record shows the original response, the rule that triggered the redaction, and the redacted response as delivered. The auditor reading the record can reconstruct the original content, the enforcement action, and the delivered content. The reconstruction is the artifact the Article 12 record-keeping mandate expects.