FDA AI/ML Guidance Audit Evidence for Device Software Reviews
FDA AI/ML guidance audit evidence should let a reviewer trace an AI-enabled device software function through its intended use, development data, validation, approved change route, deployed configuration, and postmarket signals. This guide builds a reviewer-selected evidence package while separating nonbinding FDA guidance from the binding Quality Management System Regulation.

A reviewer opens a complaint record for a missed finding and sees model release 47 in the device telemetry. The validation packet names release 46, which creates an immediate discrepancy. FDA AI/ML guidance AI audit evidence has to resolve that mismatch through controlled records, rather than a screenshot assembled after the inquiry. I would freeze deployment until the manufacturer can connect the field event to its approved change path, because uncertainty about the active model is an evidence failure with a patient-risk consequence.
TL;DR
- Start with a frozen population of AI functions and releases. Include field events, and let the reviewer select samples from that population.
- Bind each sample to intended use and risk analysis. Include development data, validation, approval, deployed configuration, and postmarket disposition.
- Label the January 2025 lifecycle document as draft guidance. Treat the August 2025 PCCP guidance as final recommendations. Treat the QMSR as a binding rule.
- Test integrity and historical retrieval. A retention setting or dashboard image cannot prove that an old record remains complete.
Freeze the evidence population before sampling
Define the population around the review question. For a release review, list every AI-enabled device software function and every production version in the period. For a complaint investigation, list the related field events and service actions. Include version changes and rejected releases. Include failed tests because they show how the control operated under pressure.
The FDA's January 2025 lifecycle and marketing submission draft guidance covers documentation intended to support FDA's safety and effectiveness evaluation and proposes a total product lifecycle risk approach. FDA marks the document as draft, nonbinding, and not for implementation. Use it as current proposed agency thinking, with that status written on the evidence index.
Freeze the manifest before sample selection. Include stable identifiers and the function, intended use, release, regulatory pathway, date, owner, and source-record location. Reviewer-selected rows give the package a defensible denominator. AI data lineage for audit can supply the join design for training and evaluation sources.
Bind each selected release to controlled records
A release sample should open into a chain of records rather than a folder of nearby files. Start with the intended use and approved claims. Link the model description and data provenance. Add the risk analysis, verification, validation, human factors work, cybersecurity assessment, approval, and deployed configuration. Add the monitoring plan and quality-system owner.
The active endpoint matters for hosted dependencies. Capture the provider and model identifier. Record the deployment region and parameter or prompt configuration under change control. Record the activation timestamp when those facts affect device behavior. An alias such as production-latest needs a contemporaneous resolution record. A screenshot with a green status dot is a concrete visual aid, but it leaves the executed configuration unresolved.
Use one correlation identifier across the release record and deployment event, then carry it into field telemetry. Test the join with a production event, then reconstruct the released function without asking the original engineer to remember where each file lives.
Treat PCCP evidence as a bounded change story
The FDA's August 2025 final PCCP guidance recommends three elements: planned device modifications; the methodology used to develop, validate, and implement them; and an assessment of their impact. FDA states that it reviews a PCCP as part of a 510(k), De Novo, or PMA marketing submission for covered AI-enabled devices.
For a sampled modification, identify the authorized plan and the planned modification. Link the executed protocol steps and acceptance results. Preserve the impact assessment and approver. Add the deployment event and post-release monitoring. If the change falls outside the authorized scope, retain the regulatory assessment and resulting submission decision.
My preference is a one-page change cover sheet with live links into controlled systems. Copying source files into an audit folder creates stale twins. A signed index preserves the package while each approved record stays under its established custody.
Prove integrity and historical retrieval
Record custody needs an identified writer and store. Document the access model and retention rule. Document the integrity mechanism. Test those properties by copying a staged record and altering the model-version field. Retain the failed integrity result. Attempt deletion with an application administrator role and keep the rejection event. Those exercises are implementation evidence, rather than language attributed to FDA guidance.
Historical retrieval deserves its own test. Ask the reviewer to select an old release and return its index and validation result. Retrieve the policy state and deployment event. Retrieve the postmarket records. Save the query and elapsed time. If archive restoration is required, preserve the procedure and restored count. The audit log chain of custody guide explains the separate write path that supports this test.
The Quality Management System Regulation final rule took effect on February 2, 2026. Revised 21 CFR Part 820 incorporates ISO 13485:2016 by reference and adds FDA requirements. Section 820.35 addresses control of records, including complaint records. The QMSR gives these records regulatory force; the two FDA AI guidance documents remain recommendations with their stated status.
Connect postmarket signals to disposition
For one complaint, trace intake to evaluation and investigation. Include the device and software version with the relevant model events. Add the clinical or engineering review, reportability assessment, correction, and closure. If the signal triggered corrective action or a proposed model modification, follow it into the controlled release path.
The FDA page on Good Machine Learning Practice guiding principles points to the January 2025 IMDRF final document and its total product lifecycle framing. Treat those principles as guidance, then use the manufacturer's QMS procedures for the actual pass criteria and records.
A request-level gateway contributes only when the device or a manufacturer-controlled application calls an LLM over the routed HTTP path. Embedded inference and local analysis need evidence from their own owners. The same applies to clinical validation and complaint conclusions. Regulatory submissions also remain with their own owners.
DeepInspect
DeepInspect sits on manufacturer-routed HTTP traffic between an authenticated application and an LLM endpoint. It evaluates application-supplied identity and workflow context. It then applies destination and content policy, inspects the response, and writes a signed, tamper-evident decision record with the policy version and timestamp.
That record can join an FDA evidence package for the requests that cross this boundary. Identity proofing and embedded inference remain with their qualified owners. Device validation and PCCP execution remain with their qualified owners. Complaint investigation, QMS decisions, and regulatory submissions also remain with their qualified owners. Book a technical deep dive at deepinspect.ai.
Frequently asked questions
- What belongs in the FDA AI audit evidence index?
Include the function and intended use with the applicable submission and PCCP. Add the released model and dependencies, risk file, validation evidence, approval, deployment record, monitoring plan, field events, and dispositions. Add source locations and record owners so the index can be tested.
- Should failed validation runs appear?
Yes. Failed runs show that acceptance rules were applied and that release control stopped an unsuitable configuration. Preserve the failure and investigation. Link the corrective work and retest. Link the final approval as part of the same chain.
- Does a PCCP authorize every future model update?
A PCCP supports modifications described within its reviewed scope and executed through its associated methodology. A change beyond that boundary needs a documented regulatory assessment and may require another marketing submission.
- Can a request log prove device safety?
A request log can prove bounded facts about routed HTTP traffic, including the supplied identity and destination. It can also record the policy version and classification. The action can be recorded with them. Safety and effectiveness conclusions require the device's validation, clinical, risk-management, and postmarket evidence.
- How should sensitive prompts appear in an assessor package?
Use access-controlled source records and a minimization method approved by privacy and quality owners. A content fingerprint can support correlation and integrity, while authorized reviewers retain a controlled route to the underlying record when inspection is necessary.