← Blog

FDA GenAI Medical Devices Paper Opens the Lifecycle Evidence Question

Parminder Singh
Parminder Singh··7 min read
Summarize with AI

FDA CDRH has opened a discussion about risk assessment, premarket evaluation, and postmarket monitoring for GenAI-enabled medical devices. The paper creates no guidance or policy change. It does expose the lifecycle evidence problem: manufacturers need records that connect the deployed device configuration and real-world interaction to clinical evaluation, monitoring, change control, and safety decisions.

Industry Verticalsai-complianceai-governanceregulationauditllmpolicy-enforcement
FDA GenAI Medical Devices Paper Opens the Lifecycle Evidence Question

FDA's Center for Devices and Radiological Health has placed a 30-page discussion paper between a fixed premarket test plan and the messy reality of a GenAI device in use. The paper covers risk assessment and competency-based premarket evaluation. It also covers clinical confirmation and postmarket monitoring. Third-party foundation models and agentic functions are included. Its status comes first: this is discussion-only. It carries neither draft nor final guidance. It also establishes no proposed or final regulatory expectations and no policy change. I think that restraint makes the paper more useful, because it exposes the evidence questions before a compliance checklist hardens around them.

TL;DR

  • FDA CDRH's September 2026 paper requests feedback. It establishes no new GenAI medical-device rule or guidance. It establishes no regulatory expectation.
  • Open-ended inputs and variable outputs make exhaustive premarket testing impractical. Evolving prompts and retrieval add further combinations. So do guardrails and foundation models.
  • Manufacturers should design a lifecycle evidence chain that joins deployed configuration and interaction records to benchmarking and clinical confirmation. The same chain should connect monitoring and change review to safety decisions.
  • An HTTP gateway can record routed policy decisions. Clinical validation and device performance monitoring stay with the manufacturer. Safety, effectiveness, and regulatory submissions stay there too.

The paper asks questions instead of setting requirements

The FDA discussion paper says its purpose is early input and broader stakeholder discussion. CDRH expressly avoids proposing a policy change or communicating regulatory expectations. It also avoids deciding whether the approaches fit existing legal authority. The paper further avoids deciding whether new authority would be necessary. Those qualifiers belong in every board memo and product ticket derived from it.

CDRH focuses on products containing one or more GenAI-enabled device software functions. FDA regulates medical devices, including qualifying GenAI-enabled devices, rather than GenAI as a technology category. The function and intended use still control the scope analysis. A clinical documentation assistant and a patient triage function can land in different regulatory positions even when they share a foundation model. So can a system that initiates an order.

The reported feedback deadline is October 19, 2026, according to The HIPAA Journal's September 11 report. A submission should answer the paper's questions with evidence and operational constraints, rather than treating the paper as an announced rule.

GenAI changes the unit of evaluation

Traditional software testing can enumerate bounded inputs and expected outputs. CDRH describes a different problem. GenAI-enabled devices may accept open-ended inputs and perform several subtasks. They may also return variable outputs for similar inputs. Behavior can also change when the underlying model or prompt changes. A change to the retrieval strategy or guardrail can have the same effect. So can a change to orchestration logic or the interface.

That means the foundation-model name alone has little evidentiary value. The paper's possible competency-based approach evaluates the final user-facing device in its deployed or representative configuration. CDRH describes device benchmarking for clinical knowledge and analytic capability. It also covers safety behavior and communication, plus generalizability. Clinical confirmation then addresses performance in real or clinically representative use.

A test report should therefore identify the complete evaluated configuration. That includes the device release and model and endpoint. It should identify the system prompt and retrieval corpus. It should also identify orchestration and safety policy, along with the user interface and relevant operating environment. AI model inventory management supplies the asset layer. The lifecycle record must also show which configuration actually handled a selected interaction.

Premarket evidence needs an explicit baseline

CDRH is considering benchmarking methods tailored to intended use and risk. The paper discusses prespecified methods and scoring rubrics. It also discusses qualified adjudicators and acceptance criteria. It also notes contamination and saturation as limits in public benchmarks. Weak representation of real-world conditions is another limit. Those are proposed discussion elements, not an FDA-prescribed test protocol.

The baseline should still be concrete. Freeze the benchmark assets and dataset versions. Record the scoring rubric and clinical source behind it. Identify adjudicators and their independence. Preserve actual outputs and disagreements. Preserve exclusions and the signed acceptance decision as well. For clinical confirmation, connect the chosen method to the device's risk and intended population. The paper discusses retrospective evaluation and shadow deployment as possible approaches. Standardized patient interactions and clinician adjudication are also included. Prospective studies offer another approach with different rigor and patient exposure.

Picture the evidence review: a clinician has interaction GC-0418 open on one screen while the release manifest occupies the other. The adjudication rubric and signed result are open beside it. A reviewer should reproduce the configuration and acceptance path without asking which undocumented prompt was active that week.

Postmarket records must connect interaction and configuration to outcomes

The paper's postmarket section considers periodic re-benchmarking and sample-based clinician review. It also considers performance-degradation monitoring. CDRH asks whether greater postmarket reliance could support greater premarket uncertainty under some conditions. It has reached no decision on that trade. Manufacturers remain responsible for postmarket monitoring of their devices under the discussion presented.

A useful operating record starts with a defined population. For each monitored period, preserve the production interaction population and sampling method. Record the active configuration and prospective review criteria. Preserve reviewer identity and result. Record the escalation and resulting governance decision. Monitoring should detect changes in the input population and data environment. It should also detect changes in underlying model components. The HIPAA AI audit trail guide shows how to keep protected interaction evidence attributable while preserving access controls.

A request record can identify a workflow and routed model call. Clinical performance evidence needs more. It connects that event to the full conversation where relevant and the device output. It also connects the clinician or patient context allowed by the protocol to the downstream disposition. Any complaint or safety signal must connect to the adjudicated outcome. Retention and secondary use require manufacturer-approved privacy controls.

Third-party model changes need their own evidence path

CDRH identifies discrete sponsor changes and passive or incremental model evolution. It also identifies unplanned changes initiated by a third-party foundation-model developer. Each can affect safety or effectiveness. The paper asks how manufacturers should detect and evaluate provider-initiated changes. It also asks how contracts and technical mechanisms might help, or how a predetermined change control plan might help.

Build the detection path before the provider release arrives. Contract terms should address model identifiers and change notices. They should cover affected capabilities and safety behavior. Evaluation information and rollback options also belong in the terms. Technical records should resolve a stable provider alias to the version observed at the request time where the provider exposes it. Change control then decides the impact assessment and re-benchmarking scope. It also decides the clinical review and deployment disposition. Post-release monitoring follows that decision.

The tamper-evident audit log guide covers integrity for event records. Integrity proves the retained record stayed intact. It offers no clinical conclusion about a changed model. That conclusion belongs to the manufacturer's qualified clinical and quality functions. Engineering and regulatory functions should also be involved.

A lifecycle evidence chain can be designed now

The discussion paper creates no required schema. Manufacturers can still prepare a chain that survives the questions CDRH has raised. I would use six joined records:

  • Approved use: Record the device function and intended user and population. Add the clinical setting and output role. State the reliance limit and accountable owner.
  • Evaluated configuration: Record the device release and foundation model. Preserve the prompt and retrieval versions. Add the orchestration and interface. Identify the benchmark assets and clinical-confirmation method. Preserve the acceptance decision.
  • Production interaction: Preserve the correlation ID and UTC time. Record the supplied user or agent context and workflow. Add the endpoint and model identifier when available. Preserve the policy decision and protected input/output references.
  • Monitoring review: Identify the source population and sample method. Record the criteria and qualified reviewer. Preserve the result and any subgroup or trajectory finding. Add the escalation.
  • Change decision: Record the detected change and impact assessment. Define the testing scope and preserve the approval. Record the deployment or rollback. Add the monitoring trigger.
  • Safety disposition: Record the complaint or signal and investigation. Preserve the clinical assessment and reportability decision. Add the corrective action and effectiveness review. Record the closure.

Use AI audit log chain of custody for the correlation and custody pattern. The joins matter more than a dashboard. A selected interaction should lead to the exact configuration and applicable evaluation. It should also lead to the monitoring result and safety or change decision.

DeepInspect

DeepInspect sits on HTTP AI traffic that a manufacturer deliberately routes between authenticated applications or agents and LLM endpoints. It evaluates application-supplied identity and workflow context against versioned content and destination policy. It also inspects the response and writes a signed, tamper-evident decision record with the active policy version and timestamp.

Those records can identify the request path and calling workflow. They can also identify the routed endpoint and policy decision. The protected interaction reference can join those details to the lifecycle chain described above. The manufacturer retains responsibility for identity proofing and route completeness. It also owns clinical validation and benchmarking, plus clinical confirmation and device performance monitoring. Safety and effectiveness remain its responsibility, along with QMS decisions and complaint handling. The manufacturer also owns FDA submissions. Local inference and direct browser activity require evidence from their actual control owners. So do embedded vendor AI and bypass traffic. Book a demo today.

Frequently asked questions

Did FDA announce new regulation for GenAI-enabled medical devices?

No. CDRH issued a discussion paper and request for feedback. It communicates no draft or final guidance and no proposed or final regulatory expectation. It makes no policy change. The paper asks about possible methods and approaches, including whether new legal authority might be needed. Existing statutes and regulations continue to govern device manufacturers. So do authorizations and applicable guidance.

Does the paper apply to every healthcare LLM?

The paper focuses on products with one or more GenAI-enabled device software functions. FDA's function-by-function, intended-use analysis remains relevant. Some healthcare software functions fall outside the device definition or another part of FDA's active oversight. Manufacturers should document the classification rationale rather than infer device status from the words generative AI.

Does FDA favor competency-based evaluation?

CDRH presents competency-based evaluation as a possible framework for discussion. The paper pairs non-clinical device benchmarking with clinical confirmation and asks stakeholders whether the approach is useful. It prescribes no benchmark suite or clinical method. It also prescribes no acceptance threshold. Any current submission strategy should rely on applicable law and current FDA guidance. Direct agency interaction should also inform it.

What does total product lifecycle evidence add?

It joins the premarket baseline to the device actually deployed. It then follows production performance and changes. Complaints and corrective decisions remain connected to the same baseline. FDA's Good Machine Learning Practice page points to 10 IMDRF principles designed around high-quality AI/ML devices that are safe and effective throughout the total product lifecycle. The September paper sharpens the open questions for GenAI's variable behavior.

Can a gateway provide the clinical evidence described in the paper?

A gateway contributes request-path evidence only when an authenticated application or agent sends HTTP traffic through it to an LLM. It can record supplied identity and workflow context. It can record content and destination policy. It can also record the model endpoint and policy version, plus the decision and response handling. Clinical validation and representative benchmarking require manufacturer-owned evidence and qualified reviewers. Clinical confirmation and whole-device safety and effectiveness require the same. Postmarket performance analysis and regulatory decisions also remain with the manufacturer.