SR 11-7 AI Audit Evidence After SR 26-2
SR 11-7 AI audit evidence needs a current-source correction before fieldwork begins. SR 26-2 superseded the 2011 letter on April 17, 2026 and excludes generative and agentic AI from direct scope. Banks that apply its model-risk disciplines to LLMs through internal policy should freeze complete populations, let reviewers select samples, trace each request to approval and change records, and preserve gaps through retest.

On April 17, 2026, the Federal Reserve replaced SR 11-7 with SR 26-2. An audit binder headed SR 11-7 AI evidence now begins with a source problem. The current guidance expressly excludes generative and agentic AI from direct scope, while a bank may still apply model-risk disciplines to those systems through its own policy. I want to focus on the fieldwork this creates: establish the governing authority and freeze the complete populations. Let the reviewer choose samples, then trace each selected LLM call through approval and policy records. Continue that trace through change and provider records to the business records.
TL;DR
- Start with SR 26-2. It superseded SR 11-7 and excludes generative and agentic AI from direct scope.
- Record the bank policy or other authority that applies model-risk disciplines to each LLM use case.
- Freeze use-case and request populations before sampling. Do the same for change and exception populations, then for incident and provider populations.
- Let the reviewer select samples, then preserve failed joins and remediation beside passing evidence. Preserve the independent retest there too.
Current authority belongs at the front of the workpapers
The Federal Reserve's SR 26-2 supervisory letter says the revised guidance supersedes SR 11-7, issued April 4, 2011, and SR 21-8. The date and applicability statement belong in the audit planning memo. The replacement chain belongs there too.
The attached Revised Guidance on Model Risk Management makes the next distinction explicit. Generative and agentic AI models sit outside its scope because those technologies are novel and rapidly evolving. The document covers conventional models, including predictive AI and non-generative, non-agentic AI.
A bank can still decide that its AI governance policy will apply selected model-risk practices to an LLM. The audit file should cite that approved internal decision and its effective date. The banking AI model-risk overview explains the source transition. Evidence work starts one level lower, with the exact policy that governs the selected system.
Build a control-to-evidence index before requesting samples
Create one index for the in-scope review period. Each row names the control objective and accountable owner. It also identifies the evidence owner and source system, the population query and retention period, and the test procedure and latest result. Add the open issue and retest reference. The source column should separate SR 26-2 principles from the bank's implementation choice.
For an LLM summarization workflow, the index may point to use-case approval in the governance repository and route records in the policy service. It may also connect releases in the deployment platform with entitlements in IAM. Provider diligence belongs in the third-party system, while outcomes belong in the case application. A single screenshot proves the screen existed on capture day. It gives the reviewer no denominator and no historical decision state.
My opinion is that a 90-page binder without a population query is weaker than a six-page package with reproducible joins. A polished binder may look finished, while the reproducible package lets an independent reviewer challenge what management selected and rerun the test.
Freeze six populations and reconcile their edges
Freeze the governed-use population first. Record every AI use case subject to the bank's policy. For each use case, capture the owner and purpose, then the model or service and risk tier. Record its status and approval, followed by its review date. Preserve the extraction query and UTC timestamp. Keep the row count and schema version, along with the operator and a content hash.
Add five connected populations for the same period:
- Routed LLM requests. Capture the correlation ID and application-supplied identity, then the purpose and endpoint. Record the policy version and classification, followed by the outcome and timestamp.
- Model and prompt changes, plus policy and route changes. Include approval and test records, release and rollback records, and closure records.
- Exceptions and incidents. Include original due dates and containment, plus extensions and final disposition.
- Identity and privileged-access events for the applications and administrators touching the control path.
- Provider usage and invoices, plus performance reports and incidents. Include approved service records.
Reconcile the edges before samples are chosen. Provider usage should align with approved routes and request totals at an explained level of granularity. Deployment records should account for every policy version observed. An unmatched provider event is a bypass indicator. An unmatched request may expose a telemetry or billing difference. Preserve both until investigation closes them.
Reviewer-selected samples need complete traceability
Give internal audit or another independent reviewer the frozen manifests and data dictionary. The reviewer should select allowed and denied requests. The selection should also include one escalation and an event under a changed policy, plus an exception and any bypass indicator. Management can explain the records after selection rather than replacing awkward cases with cleaner examples.
For each selected call, trace the use-case approval and current owner. Join the caller identity supplied by the application to the relevant entitlement snapshot. Confirm the destination and content class, then retrieve the exact policy version and decision. Connect the event to response handling and the business application's disposition. A route log proves the policy point's action; the case system proves what the bank did with the model output.
Set two artifacts side by side on the review screen: the CSV request row and its release ticket. The route and policy version should agree with the release identifier. AI audit log chain of custody provides the export and custody pattern once those records enter the examination workroom.
Test design and operation separately
SR 26-2 retains risk-based principles around effective challenge and validation. It also addresses ongoing monitoring and governance, inventory and documentation, and third-party products for models inside scope. When the bank maps those disciplines to an LLM through policy, audit should state the mapping as a bank control rather than language imposed directly on generative AI by the supervisory letter.
Design testing asks if the control could achieve its stated objective. Inspect role definitions and approval criteria, then policy logic and restrictive failure behavior. Review evidence fields and retention, followed by escalation paths. Operating testing asks if the control ran across the period. Use the frozen population and reviewer-selected samples. Add controlled failures and access history, then test historical retrieval.
Run a non-production test with missing identity context and an unapproved model route. Include an expired exception in the same test. Record expected outcomes before execution. Keep raw results, tester identity, timestamps, and defects. A current configuration screenshot should never erase the failed first run.
Third-party evidence needs a bank-owned join
The 2026 model-risk guidance discusses vendor and other third-party products for models inside its scope. LLM governance also intersects with the agencies' Interagency Guidance on Third-Party Relationships, which applies to banking organizations' third-party relationships under a risk-based approach.
For each selected provider, join due diligence and contract terms to the service the bank actually approved. Preserve service descriptions and data-use terms. Keep subcontractor information and incident notices, along with performance reports, exit provisions, and the complete issue history. Provider documentation can describe the service. Bank records must prove which route and use case entered production.
A provider assurance report leaves several audit questions open. It says little about the bank's supplied identity context and route completeness. It also leaves questions about business use of output and internal exceptions, as well as outcome monitoring. Third-party AI risk management covers the embedded-model problem that appears when inference remains inside vendor-controlled software.
Failed evidence stays in the package through retest
Classify each finding by the objective it affects. Categories include scope and population completeness, authorization and change control, monitoring and third party, and retention and integrity. Business outcome is another category. Keep the original observation and due date. Add accountable ownership and interim containment. Record the root cause and remediation artifact, followed by independent retest and closure authority.
A missing join deserves its own finding. If a routed request has no approved use case, or a provider event lacks a request record, record the gap before debating severity. The audit trail should retain the first failed result beside the later pass. Replacing it makes the file look cleaner and destroys the remediation history an examiner may ask to inspect.
Historical retrieval belongs in the closing test. Ask the evidence owner to reproduce a selected decision under the policy version active on that date, then verify integrity after export. Tamper-evident audit logs for AI explains the separate write-path property, but the bank still owns retention and legal hold. It also owns access approval and production to supervisors.
DeepInspect
DeepInspect supports the routed HTTP portion of a bank's internally defined AI control environment. It sits between authenticated users or agents and approved LLM endpoints. It evaluates application-supplied identity and workflow context against content and destination policy, then records the active policy version and decision before an allowed request proceeds.
Each routed decision produces a signed, tamper-evident record outside the calling application's write path. Those records can support request populations and reviewer samples, plus controlled-failure tests and historical retrieval. The bank retains responsibility for identity proofing and route completeness. It also owns model governance and validation, outcomes analysis and third-party oversight, and retention and independent audit. Direct browser use and local inference need evidence from their actual owners. Vendor-managed model calls do too. Book a technical deep dive at deepinspect.ai.
Frequently asked questions
- Is SR 11-7 still the current Federal Reserve guidance?
SR 26-2 superseded and replaced SR 11-7 on April 17, 2026. Current audit plans should cite SR 26-2 and preserve SR 11-7 only as historical program context. Existing policies and procedures carrying the old identifier need a source review. Workpapers carrying that identifier need the same review, with any continuing internal requirement mapped to its current authority.
- Does SR 26-2 directly cover generative AI?
The revised guidance expressly places generative and agentic AI outside its scope. A bank may still apply its governance disciplines to those systems through internal policy or another applicable authority. The audit report should name that source and avoid presenting an internal mapping as a direct SR 26-2 requirement.
- Can a request log prove model validation?
A routed request record can prove the context and policy decision observed at the HTTP boundary. Model validation examines suitability and reliability, plus limitations and performance under the bank's approved methodology. Validation reports and test data sit with model-risk and business owners. So do challenge records and outcomes analysis, along with monitoring evidence.
- Who should choose the samples?
Internal audit or an independent validation function should select samples from frozen populations. Another qualified reviewer may also do so. Control owners can supply context and investigate differences. Preserve the selection method and original denominator, along with failed joins and replacements requested by neither management nor the system owner.