OCC AI Audit Evidence for Model-Risk Governance
OCC model risk AI audit evidence should connect a bank-approved LLM use case to the routed request population and reviewer-selected samples, then connect model and policy changes to monitoring results, exceptions, and third-party oversight. This guide reflects OCC Bulletin 2026-13, which replaced the 2011 model-risk guidance and excludes generative and agentic AI from its direct scope.

An OCC examiner selects routed event LLM-48217 from a quarter-end export. The row shows an authenticated fraud analyst and a permitted provider endpoint under policy version 18. The bank still has to produce the approved use case, show what version changed, identify the monitoring result, and trace any model output used in the investigation. OCC model risk AI audit evidence begins with that join. I would send back a polished binder that lacks the source records. Examiner-ready evidence lets the reviewer choose a row and follow it without a guided tour from the developer who built the integration.
TL;DR
- Start with OCC Bulletin 2026-13. It replaced OCC Bulletin 2011-12 and excludes generative and agentic AI from direct scope.
- Freeze and reconcile four population groups before sampling: governed use; routed requests; changes and exceptions; providers.
- Trace reviewer-selected calls through approval and policy, then through response handling, monitoring, and business disposition.
- Treat gateway records as request-decision evidence. Model validation and governance remain with the bank's assigned functions.
Fix the supervisory source boundary first
On April 17, 2026, OCC Bulletin 2026-13 rescinded OCC Bulletin 2011-12. The Federal Reserve's matching SR 26-2 letter superseded SR 11-7. The revised interagency guidance covers model development and use, validation and monitoring, governance and controls, plus vendor products. It uses a risk-based approach and says it sets neither enforceable standards nor prescriptive requirements.
The scope line matters for this article. Bulletin 2026-13 expressly places generative and agentic AI outside the revised guidance. It also says a banking organization's governance practices should determine appropriate controls for tools and systems outside the document. A bank may therefore apply familiar model-risk disciplines to an LLM under its own policy, based on purpose and exposure, but the evidence package should identify that as the bank's governance decision.
There is no OCC-specific AI evidence checklist hiding behind the old bulletin number. The first workpaper should name the current source and the exact control scope, then record the bank policy that brings the use into review.
Build a control-to-evidence index
Start with one index that exposes the chain an examiner will test. Each line names the bank control and accountable owner; applicable use cases and source system; record type and retention rule; population query and latest test; open issue and evidence location. Put the policy statement beside the operating record rather than scattering them across folders.
For an authenticated HTTP LLM workflow, the index should cover five evidence groups:
- Governed use cases: approved purpose and business process; owner and model or provider route; permitted data and reliance level; human review and limitations; monitoring plan and reassessment triggers.
- Routed decisions: correlation ID and UTC timestamp; application-supplied principal and workflow context; endpoint and model identifier when available; content classification and policy version; decision and response disposition.
- Changes: model and endpoint changes; system prompt and retrieval source changes; policy and connector changes; application and monitoring-threshold changes, including emergency work and rollback.
- Exceptions and issues: denied requests and approved overrides; bypass findings and missing identity; monitoring breaches and incidents; complaints and remediation; retests and closure authority.
- Providers: due diligence and contract; approved services and data handling; assurance reports and incidents; performance reviews and subcontractor information where appropriate; exit planning.
AI model inventory management covers the broader asset record. This index focuses on the evidence joins used during OCC fieldwork.
Freeze populations and reconcile the routes
Export the five populations for the examination period before anyone selects samples. Preserve each query and filter with its extraction timestamp and row count. Preserve the schema version and content hash too. Record the person who performed each export. A CSV open on the left monitor and a change ticket open on the right should show the same model route and release identifier.
Reconcile the routed-decision export against records outside the policy gateway. Provider usage or billing should match approved destinations. Compare application egress/API-management records with calls sent by in-scope applications. Deployment records should account for every active endpoint and policy version. IAM records should explain the principals and service identities observed during the period.
Document every difference and assign its investigation. Direct provider calls expose a bypass path. An endpoint in billing but absent from the approved-use register may reveal an ungoverned use. A decision row with a shared service identity may prove the technical caller while leaving the person or agent unresolved. The post-authentication gap explains why the application must supply that acting identity and workflow context.
Sampling starts after the denominator survives these joins. A management-selected folder of successful requests proves very little about population completeness.
Reconstruct one reviewer-selected call
Give the examiner the frozen manifest and data dictionary, then preserve the selection method. Include a permitted call and a denial. Add two change-related samples: an approved exception and a request after a material change. Add a call tied to a monitoring or customer-impact event when the period contains one.
For each selected row, begin with the bank's use-case approval. Retrieve its purpose and owner; permitted data and reliance limit; provider route and test plan; active exceptions. Join the row to the authenticated principal or agent and the business workflow. Preserve the input reference or protected source location and model endpoint; policy version and policy decision; response reference and application handling. Finish with the human review or downstream business disposition that shows how the output was used.
That last join separates request telemetry from model-risk evidence. A gateway can show that version 18 permitted a fraud analyst to send classified case text to an approved endpoint at 14:07 UTC. The bank's fraud and model-risk records, together with its internal audit records, must show the output review and case decision, plus the monitoring result and any corrective action. AI audit log chain of custody provides the custody pattern for keeping those records connected.
Test operating evidence under controlled failure
A current policy page demonstrates design. Operating effectiveness needs dated tests and selected production records. Write the expected result before each test and retain four pairs: input and environment; tester and time; raw output and actual result; issue and retest.
Use a controlled test environment to submit an approved request, then try an unapproved provider destination. Remove required identity context. Attempt a restricted data class under a lower-privilege role. Attempt to invoke an expired exception. Retrieve an event near the start of the bank's retention period. Modify a staged copy and run the documented integrity check. Test an application route that could bypass the gateway and preserve the network or provider evidence used to detect it.
The 2026 revised guidance describes internal audit's role as evaluating the rigor and effectiveness of model-risk practices and policy implementation when audit participates in the program. Keep four roles visible in the package: control owner, model-risk reviewer, test operator, and internal auditor. Independence belongs in the workpaper, not in a label applied after a failed test.
Connect changes, monitoring, and outcomes
An LLM evidence package needs a dated history. Preserve the approved baseline, then link five change classes to the record: model; prompt and retrieval; policy; endpoint. Connect each change to impact assessment and testing, approval and deployment, monitoring and rollback, closure. If the provider changes a model behind a stable API name, record when the bank learned of it and which use cases were reassessed.
Monitoring should match the approved purpose. A summarization use may review unsupported statements and prohibited data handling. A classification use can compare selected outputs with adjudicated labels. A fraud-support workflow should connect sampled model output to analyst disposition and later case information. Record four pairs: thresholds and review frequency; population and sample method; result and escalation; the governance decision that followed.
The result belongs beside open recommendations and exceptions. Keep the original due date when work slips. Preserve risk acceptance and compensating controls, then the independent retest and closure approval. The current interagency model-risk guidance says documentation supports two related pairs: recommendations and responses; exceptions and remediation. That gives an examiner a continuous history instead of a quarter-end reconstruction.
Add third-party evidence without outsourcing ownership
The current OCC third-party risk bulletin applies a continuous risk-management life cycle to bank relationships and scales oversight to the risk and criticality involved. Its underlying interagency guidance identifies four documentation groups: inventories and risk assessments; due diligence and contracts; performance, incident, and remediation records; independent reviews and board reporting. It also says examiners may perform transaction testing or review testing results.
For an LLM provider, connect those artifacts to the bank's approved route and actual production population. Show the endpoint covered by the contract and how provider changes reach the bank. Show the applicable data handling and remaining assurance gaps. Show how incidents or deteriorating performance trigger action. Preserve any limit on audit rights as a risk decision rather than smoothing it out of the file.
Provider documentation supports oversight. It cannot establish that the bank used the service only for approved purposes or that a selected output met the bank's acceptance criteria. Those conclusions require bank-owned route reconciliation and testing, backed by monitoring and outcome records.
DeepInspect
DeepInspect sits on authenticated HTTP traffic that a bank routes between its applications or agents and LLM endpoints. It evaluates application-supplied identity and workflow context against two policy inputs: content and destination. It inspects the response and writes a signed, tamper-evident decision record containing the active policy version and timestamp.
Those records can supply the routed-decision population and selected request evidence described here. The bank retains responsibility for three control pairs: identity proofing and route completeness; model governance and validation; business outcome review and provider oversight. It also retains responsibility for retention decisions and independent audit. Local inference and direct browser use need evidence from their actual control owners. The same applies to embedded vendor AI and any bypass path. Book a technical deep dive at deepinspect.ai.
Frequently asked questions
- Is OCC Bulletin 2011-12 still the current model-risk guidance?
OCC Bulletin 2026-13 rescinded Bulletin 2011-12 on April 17, 2026. The Federal Reserve issued the corresponding SR 26-2 letter, which superseded SR 11-7. A current workpaper should cite the 2026 source and use the 2011 material only as historical program context. The revised guidance also excludes generative and agentic AI from direct scope, so the bank should preserve its own policy decision for applying model-risk governance to an LLM use case.
- Does OCC have an AI-specific audit-evidence rule?
The cited OCC model-risk bulletin creates no AI-specific evidence rule. Bulletin 2026-13 describes nonprescriptive guidance and expressly excludes generative and agentic AI. Examiner-ready artifacts described here are a recommended evidence pattern for a bank that governs LLM use through its own risk program. Applicable laws and bank policy may create separate record or control duties, as may the business use and other supervisory guidance.
- Who should select the request samples?
An examiner or an independent review function should select rows from the frozen population. Control owners can explain systems and investigate exceptions, but they should preserve the original denominator and failed samples. Document the sampling method and include four targeted item types beside any random selection: denials, changes, exceptions, and bypass indicators.
- Can a content fingerprint replace the prompt and response?
A fingerprint can support correlation and integrity while reducing duplicate storage of sensitive bank information. Authorized reviewers may still need access to the protected source content for model testing and investigation, or for outcome review. Document four pairs: normalization and covered fields; key custody and source location; retention and access approval; the process and result used to reproduce the integrity check.
- Does an HTTP gateway validate an LLM?
Model validation stays with the bank's assigned model-risk or independent review function. A routed gateway can contribute request-level evidence across three pairs: supplied identity and destination; classification and policy version; decision and response handling. It offers no conceptual-soundness conclusion or output-accuracy finding. It also offers no outcomes analysis or coverage for calls that bypass the route. Those elements need separate owners and records.