← Blog

CSBS AI Supervisory Framework: Prepare the Examiner Evidence File

Parminder Singh
Parminder Singh··6 min read
Summarize with AI

The CSBS Artificial Intelligence Supervisory Framework gives state examiners a Core Examiner Guide, document request list, work program, nonbank supplements, and an optional risk-tiering worksheet. Financial institutions can prepare by connecting each AI use case to its owner, purpose, vendor, risk tier, data controls, change record, and operating evidence. The framework remains discretionary, and each state agency decides how to incorporate it.

Industry Verticalsai-complianceai-governancecomplianceregulationauditpolicy-enforcement
CSBS AI Supervisory Framework: Prepare the Examiner Evidence File

On September 16, 2026, the Conference of State Bank Supervisors released a practical examination toolkit for artificial intelligence. The CSBS Artificial Intelligence Supervisory Framework gives state examiners a common way to identify AI use, scope a review, request records, and decide where existing supervisory resources deserve a closer look.

The first preparation step is concrete. Pick one deployed AI use case and trace it across the institution before an examiner asks: business purpose, accountable owner, vendor or model, affected data, risk tier, approval, monitoring, and evidence of what happens in production.

TL;DR

  • CSBS released the framework as a discretionary examiner tool, and each state agency decides how much of it to use.
  • The Core Examiner Guide starts with scoping, then requests governance records, an AI inventory, use-case documentation, vendor evidence, and control testing.
  • Banks and nonbanks should prepare one traceable evidence file per material use case rather than a disconnected folder of policies.
  • DeepInspect can support the narrow HTTP LLM-traffic slice. Institution-wide governance, model validation, vendor oversight, and consumer protection remain separate work.

The framework changes the examination conversation

The CSBS framework creates no new legal obligation. Its Core Examiner Guide, version 1.0 says that directly and calls for proportional use based on an institution's size, complexity, risk profile, and AI use. Each state financial regulatory agency decides how the material enters its own supervisory program.

That discretionary status matters during preparation. A compliance team should avoid presenting September 16 as a universal effective date or treating every worksheet field as a binding requirement. The operational signal is still strong: examiners now have named scoping questions and a document list that can make a vague AI discussion specific within minutes.

The initial scope asks about AI in products, services, operations, compliance activities, and internal support. It also asks about embedded vendor AI, customer-facing activity, generative AI plus sensitive information and risk differentiation. A negative answer can be checked against vendor inventories, software inventories together with approved tools and recent platform changes. An institution that says it has no AI therefore needs evidence for the answer.

Build the document request file before fieldwork

The Core Examiner Guide identifies the materials an examiner may request. Governance records include policies and committee materials, management or board reporting, dashboards and training, alongside assigned responsibility. Inventory records include systems, tools, models, business purposes, ownership, risk assessments, vendor products with possible embedded AI, and change documentation.

Generative AI adds another set of records: approved tools, restrictions, data controls, output review, vendor oversight, testing, with internal assessments and sample consumer-facing outputs where applicable. Contract language about retention, sharing plus training use and institutional data also appears in the request list.

Put those items into a standing evidence file indexed by use-case ID. A loan-servicing assistant named AI-042 should have one row in the inventory and one folder that carries its approval, vendor review, operating owner, data classification, model route, testing result, monitoring report together with change history and contingency plan. At 8:30 on an examination morning, a labeled folder is more persuasive than six owners searching separate systems.

Use the AI governance framework as the institution-wide index, then attach each CSBS evidence folder to its exact use-case ID.

The examiner will follow one use case

The framework's inventory section asks where AI operates, what purpose it serves and who owns it, alongside how its risk tier is assigned. It covers internal tools and customer-facing systems, including decision support, automated decisions and embedded vendor capability. It also asks how the inventory stays current as use changes.

Prepare for a vertical trace rather than a policy recital. Start with one use case, then show the business process and decision influence. Identify the model or service, deployment type, vendor relationship, affected customer population, input data, output review, approved operators plus monitoring cadence and last material change. The risk tier should lead to actual review depth.

I think a beautifully formatted AI policy with no route to a deployed use case is weaker than a plain spreadsheet whose links all open. Examiners can test the spreadsheet. A policy that floats above production gives them more questions.

The AI vendor risk assessment template can support the vendor branch of this trace. Keep the institution's own outcome monitoring and approval evidence beside the vendor package, because provider assurance cannot establish that a specific deployment remains within its approved purpose.

Generative AI gets an explicit workstream

The Core Examiner Guide names large language models and agentic tools. It asks institutions to distinguish public, private together with internally deployed and vendor-provided tools. Review areas include data controls, output oversight, inaccuracy, hallucination, prompt injection, data exposure and evolving restrictions, alongside monitoring.

Systems that take actions with limited human direction receive added attention. The guide points examiners toward permission boundaries, human checkpoints, logging, with reversibility and the ability to restrict or halt a system. For an agent that drafts customer communications, preserve its permitted actions and escalation rule alongside samples of reviewed output. For an internal research assistant, record the approved data classes and provider route, then test a prohibited input and retain the result.

A useful operating test runs one permitted request and one denied request through the production-like LLM route. The evidence should show the originating identity, use-case ID, route, policy version, decision plus timestamp and response handling. That artifact supports the runtime portion of the use case. It leaves model validation together with consumer impact review and broader governance with their assigned owners.

Nonbank supplements require a separate preparation track

CSBS published Nonbank AI Supplements, version 1.0 for third-party and vendor oversight and model risk, alongside consumer protection. They supplement existing review resources and can overlap on one use case. A vendor underwriting tool, for example, may trigger all three areas.

The vendor supplement asks about embedded capability, data practices, intellectual property, model changes, incident terms, audit information, with ongoing monitoring and fallback planning. The model-risk supplement focuses on intended use, limitations, validation, drift plus human challenge and explainability. Consumer protection review covers decision influence, specific reasons, communications, fair lending, UDAP or UDAAP together with proxy effects and outcome differences.

Nonbanks should tag every material use case with the applicable supplement. The next step assigns an owner to each evidence set. Procurement may own contracts, model risk may own validation, compliance may own consumer testing, and technology may own runtime controls. One accountable use-case owner should be able to assemble the whole package without pretending one team performs every control.

A preparation exercise for the next examination

Run a 90-minute mock request with technology risk, compliance, model risk and internal audit, alongside one business owner. Give the group a single use-case ID and ask for the current inventory row, approval, risk tier, vendor terms, data-flow diagram, output review, monitoring evidence, with latest change and contingency plan.

Record missing items as control gaps with named owners and dates. Repeat the exercise for one internally developed tool and one embedded vendor capability. For institutions subject to broader operational-resilience duties, the DORA AI compliance guide for banks gives a related view of accountable operations and evidence.

The mock review should also test retrieval. A control may operate correctly while its evidence sits in a ticketing queue that the examination team cannot access. Export a stable copy, record the source system, and preserve the reviewer's sign-off. That small discipline turns operational traces into an examiner-ready package.

DeepInspect

DeepInspect covers authenticated HTTP traffic deliberately routed between users or agents and LLM endpoints. For that slice of an institution's AI inventory, it can evaluate application-supplied identity context and prompt classification against policy before a request reaches the model, then create a per-decision audit record.

That record can support a use-case file with identity, route, policy version, classification plus outcome and timestamp. It can also demonstrate an operating test for an allowed request and a denied request. The AI agent post-authentication gap explains why an authenticated service credential alone provides too little context for a per-user decision.

DeepInspect does not create the institution-wide AI inventory, validate models, approve vendors, test fair-lending outcomes, manage endpoints or protect stolen credentials. It also leaves local execution, STDIO, provider training and institution-wide CSBS readiness to other controls and owners. Direct traffic that bypasses the proxy produces no DeepInspect decision record. Let's talk today.

Frequently asked questions

Is the CSBS framework a new binding rule?

CSBS describes it as a discretionary tool, and the Core Examiner Guide says it creates no new legal obligations or supervisory requirements. Each state agency determines how it enters that agency's program. Institutions should track adoption and examiner communications in their jurisdictions while preparing the records the framework makes visible.

What should an institution produce first?

Start with a current inventory connected to real use-case evidence. Each material row should name a business purpose and owner, deployment type, vendor or internal model, affected data, risk tier, approval status and monitoring approach, alongside last change. The record should link to supporting documents rather than repeat them.

Does a vendor SOC report answer the examiner's questions?

A vendor assurance report supports third-party review. The institution still needs evidence for its selected use case, including intended purpose, data handling, output oversight, performance monitoring, change review, consumer impact where relevant, and contingency planning. The Nonbank AI Supplements explicitly direct attention to ongoing oversight rather than sole reliance on vendor representations.