FFIEC AI Audit Evidence: Build an Examiner-Ready Record
FFIEC AI audit evidence should let an examiner start from a complete population of AI systems, providers, releases, changes, incidents, and exceptions, then select samples and reconstruct what happened. This guide applies technology-neutral FFIEC examination procedures to AI without presenting them as a dedicated AI rule.

An examiner asks for every AI-related production change made during the quarter. The team supplies twelve approved tickets, but the deployment log contains fourteen model or prompt releases. That two-row difference is where FFIEC AI audit evidence starts. The institution needs a fixed population and an explanation for every difference. The records must also let the examiner choose a sample without management curating the answer first.
TL;DR
- Freeze complete populations before samples are selected. Reconcile inventories to provider and deployment records. Repeat the reconciliation against change and incident records, then access records.
- Let the reviewer choose samples, then trace each one through the risk decision and operation. Continue through monitoring and closure.
- Test historical retrieval and record integrity. A policy page shows design; a selected transaction shows operation.
- Treat this as an AI application of technology-neutral FFIEC guidance, not an FFIEC AI-specific rule.
Start with the source boundary
The FFIEC has not published a standalone AI control rule in the three IT Examination Handbook booklets used here. Its technology-risk work programs still give examiners a practical route into AI systems.
The Architecture, Infrastructure, and Operations booklet says its examination procedures may be applied according to an entity's size and complexity. Its business also shapes how the procedures apply. Depending on scope, examiners may sample a line-of-business process or review it at enterprise level. Its procedures call for inventories and current diagrams, along with policies and audit reports. They also cover testing results and third-party reports. Change records and metrics are included, as is corrective action.
The Development, Acquisition, and Maintenance booklet applies risk mitigation across the lifecycle and describes its examination work as technology-agnostic. Its August 2024 procedures cover governance and risk assessment. They also cover requirements and testing, followed by acquisition and maintenance. Activity logs and changes are included, as is follow-up. The glossary includes AI, but that inclusion does not create a separate AI regime.
The Information Security booklet adds access and logging. It also covers security monitoring and assurance, along with independent testing. Translate these sources into AI evidence while preserving their actual status and scope.
Freeze six populations before sampling
An evidence request should begin with dated exports from systems of record. Preserve the query and filters. Record the extraction time and row count. Name the owner and retain a content hash. A spreadsheet assembled from memory has no defensible denominator.
Build six linked populations for the review period:
- AI use and asset population: include production and pilot systems; internal and vendor-embedded systems; local and externally hosted systems. Record the owner and business process; data class and criticality; provider and model; status and route.
- Release and change population: include model and prompt changes; retrieval source and policy changes; endpoint and connector changes; access and infrastructure changes. Include rejected and emergency work.
- Request and decision population: include routed events with stable correlation IDs and timestamps. Record application and principal context; destination and data classification; policy version and outcome; response disposition.
- Access population: include privileged and service accounts. Record role grants and recertifications, plus break-glass activity and denied attempts for AI components and evidence stores.
- Incident and exception population: include alerts and overrides; policy exceptions and complaints; model failures and unauthorized use; investigations and risk acceptances; overdue actions.
- Third-party population: include AI providers and embedded capabilities; subprocessors and contracts; service levels and assurance reports; issues and termination dependencies.
Reconcile those exports against independent records. Provider billing can expose an endpoint missing from the request inventory. Cloud and API gateway telemetry can reveal a route absent from the architecture diagram. Deployment events can identify releases with no approved change. Accounts payable and single sign-on records can uncover an unregistered AI service.
I would stop sample testing when these joins do not reconcile. Sampling an unreliable population produces a neat binder and a weak conclusion.
Let the examiner select the samples
The AIO booklet expressly allows sampling based on examination scope. Give the reviewer the frozen population and data dictionary, then preserve the sample-selection method. Include random rows where appropriate. Target a failed test and an emergency change. Add a privileged action and a policy override. Include a customer-impacting event and a third-party outage.
For every selected AI system, retrieve its approval and risk assessment. Identify the owner and preserve the architecture and data-flow record. Add the data classification and provider review. Include the contract and access design, followed by the testing plan and acceptance criteria. Preserve the production release and monitoring records, then the latest review. The selected item should trace to source records rather than copies scattered through an audit folder.
For every selected change, connect the request to impact analysis and approval. Preserve the build artifact and test results. Add segregation of duties and implementation records, followed by the version and rollback plan. Include the post-implementation review and closure. FFIEC's development and maintenance procedures ask examiners to review change requests and risk assessments. They also review approvals and change logs, plus version controls and testing. Back-out plans and post-implementation review complete the sequence. Apply that sequence to model and prompt changes, as well as retrieval and policy changes, according to risk.
For every selected event, reconstruct the operating story. Show the active versions and supplied identity. Identify the applicable rule and data classification, followed by the provider destination and decision. Preserve the output handling and monitoring result. Include any investigation. AI audit log chain of custody describes how to preserve the joins without treating the application under review as the sole custodian.
Prove design and operating effectiveness separately
The Information Security booklet distinguishes assurance over system design from assurance over control operation. An approved standard can show design. It cannot show that a control ran on Tuesday at 14:07 UTC.
Test one expected success and one expected failure for each material control. Attempt an unapproved provider destination. Submit a request without required identity context. Change a staged record and retain the integrity alert. Use a privileged role to attempt deletion. Retrieve an archived event from the earliest retained period. Select a disabled account and verify that its token no longer reaches an AI endpoint.
State the expected result before running the test. Preserve the input and environment. Record the tester and time. Keep the actual result and raw output. Document the deviation and retest. The Information Security booklet calls for a documented testing and evaluation plan. The plan covers components and methods, plus timing and frequency. It also defines acceptance criteria. The booklet says testing and audits should have coverage and depth appropriate to the assurance sought. Independence should also match that assurance.
A self-assessment remains useful, but FFIEC says it does not eliminate independent audit. Keep the control owner and tester visible on the evidence index. Identify the approver separately.
Preserve logs as evidence, not decoration
The Information Security booklet describes logs as important to incident investigation. It recommends retention policies and restricted access. It also recommends adequate capacity and protected backups, plus periodic independent review of logging practices. Its examination procedures look for audit trails of changes and enabled logging.
For AI events, document who can write and read the records. Record who can alter and export them. Identify who can delete the records. Name the authoritative store and retention rule. Preserve time synchronization and schema versions. Monitor the evidence pipeline for collection gaps. Test whether a correlation ID can join the application event to the policy decision and provider request, then to response handling and downstream action.
Minimize sensitive content without destroying reconstructability. A content fingerprint can establish correlation and integrity, while an authorized path to the protected source record supports necessary review. Privacy and legal owners should approve that method. Records owners should approve it too. Tamper-evident audit logs for AI covers the separate write path and integrity controls.
Package open findings with their history
Examiners begin by reviewing past reports and outstanding issues. They examine management responses and root-cause correction, followed by retesting. Build an issue population that retains the original finding and risk rating. Identify the accountable owner and due date. Preserve the interim measure and status changes. Include evidence of correction and the independent retest. Record the closure authority.
Do not overwrite a missed due date with a new one. Keep both dates visible in the issue history. Repeat issues should point to the earlier finding and explain why the prior correction failed. Board and committee materials should match the source issue register, including overdue work and accepted risk.
Third-party evidence needs the same discipline. Preserve due diligence and contract controls. Keep service levels and provider reports. Record unresolved exceptions and monitoring, along with exit planning. A SOC report can inform assurance, but the AIO booklet expects management to assess provider reports. Management should monitor security and performance, plus outstanding issues and root causes. Action plans must also remain visible. Map each relevant exception to the institution's AI service and compensating control.
DeepInspect
DeepInspect sits on routed HTTP traffic between an authenticated application or agent and an LLM endpoint. It evaluates application-supplied identity and workflow context. It applies destination and content policy, inspects the response, and writes a signed, tamper-evident decision record with the policy version and timestamp.
Those records can support the request-and-decision population for traffic that crosses this boundary. IAM remains responsible for identity proofing. Application and model-risk teams retain their own decisions and records. The same applies to data and third-party teams. Legal and compliance teams remain responsible for their records, as does internal audit. Embedded and local inference require separate evidence. Book a technical deep dive at deepinspect.ai.
Frequently asked questions
- Does FFIEC require an AI-specific evidence binder?
The cited booklets set technology-risk and examination expectations. They do not prescribe a dedicated AI binder. A linked evidence index is a practical way to apply those expectations to AI systems and answer scoped examiner requests.
- How large should the sample be?
Use the examination objective and population quality. Account for risk and control frequency, then apply the reviewer methodology. Preserve the sampling rationale in the workpaper. A high-risk exception or emergency change may be selected deliberately even when it falls outside a random sample.
- Can screenshots satisfy an evidence request?
Screenshots can clarify a view at a moment in time. Source exports and query records provide stronger evidence. Approvals and raw test output add support, along with protected logs. Together, these records show completeness and custody across historical operation.
- What should happen when inventory and provider records disagree?
Open an issue and identify the missing or misclassified system. Assess exposure, then update the inventory through controlled workflow. Test the detection process. Preserve the original mismatch. Silent cleanup removes the evidence needed to evaluate the control.
- Where does an HTTP gateway fit?
It can provide evidence for authenticated application-to-LLM HTTP requests that traverse it. Embedded inference, local models, direct browser sessions, and opaque vendor-managed AI remain outside that route and need evidence from their own control owners.