← Blog

EU MDR AI Audit Evidence for a Notified Body Review

Parminder Singh
Parminder Singh··5 min read
Summarize with AI

EU MDR AI audit evidence should let a notified body move from a selected device to its intended purpose, approved software configuration, risk and validation records, field signals, changes, and corrective actions. This guide builds that assessor package while keeping routed LLM evidence inside its actual HTTP boundary.

Industry Verticalsai-complianceai-governanceauditregulationpolicy-enforcementforensic-audit
EU MDR AI Audit Evidence for a Notified Body Review

A notified-body reviewer selects one device configuration, opens its technical file, and follows a field complaint to the software version that produced the disputed output. EU MDR AI audit evidence succeeds when every join in that chain returns a controlled record. The current consolidated Medical Device Regulation requires technical documentation to remain current and contain the elements in Annexes II and III. I would build the review package around retrieval and traceability, rather than screenshots collected the week before assessment.

TL;DR

  • Freeze the device population and let the reviewer select a representative sample. Bind each row to intended purpose and risk class. Add the Basic UDI-DI, software configuration, plus approval state.
  • Cross-reference risk and verification through stable identifiers. Add clinical evaluation, post-market signals, changes, plus CAPA. Test each join.
  • Prove integrity and historical retrieval with a copied-record alteration test and an old-record query.
  • Keep routed HTTP LLM evidence separate from device validation and clinical conclusions. Notified-body sampling plus regulatory submissions also remain separate.

Start with the device and intended-purpose anchor

The package begins with the device selected under the applicable conformity-assessment procedure. MDCG 2019-11 rev.1 treats intended purpose as central to software qualification and classification. Its Rule 11 discussion connects software used for diagnostic or therapeutic decisions to the possible impact of an incorrect decision.

Create a sample manifest with the device name and Basic UDI-DI. Add intended purpose, patient population, MDR class, plus the applicable classification rule. Finish with the certificate reference and active software configuration. Add the selection rationale and the date the population was frozen. A blue folder labelled Class IIb, release 4.2 is useful only when its index resolves to controlled source records.

The reviewer should select the sample or approve the repeatable selection method. MDCG 2025-6 confirms that the governing MDR sampling rules remain applicable to high-risk medical-device AI. It supplies no universal runtime event count.

Build a cross-reference spine through Annex II

Annex II requires technical documentation that is clear and organised. It must also be readily searchable and unambiguous. The file covers intended purpose and users, patient population, operation, plus qualification rationale. It adds risk-class justification and configurations. Design information and the precise identity of controlled conformity evidence complete the index.

Turn that structure into an evidence index. Each sample row should link to the approved requirements baseline and architecture. Add the risk-management file, verification results, plus validation results. Clinical evaluation and labelling follow. Supplier records and release approval complete the row. A stable software or model identifier matters here. So do endpoint and dependency versions when a hosted component can alter device behaviour.

The AI data lineage guide provides a useful pattern for binding datasets and evaluated artifacts to a released version. The quality system still owns the regulatory meaning of those records. A runtime event supports traceability; it never establishes clinical performance by itself.

Reconcile Annex III field signals to action

Article 83 requires an active and systematic process for gathering device data throughout its lifetime. The process analyses quality, performance, safety, plus the resulting action. Annex III names serious and non-serious incidents, trend information, literature, plus databases. It adds registers and user feedback. Complaints and public information about similar devices also belong. The plan needs methods and indicators, followed by threshold values. Investigation protocols and corrective-action procedures complete it.

For the selected device, export the field-signal register and reconcile it to the approved post-market surveillance plan. Choose one complaint and one non-serious event when those records exist, then add a threshold review. Trace each signal through triage and risk assessment. Continue through investigation, disposition, plus any CAPA or change request. Record explicit none found results for empty categories and preserve the query.

Article 85 assigns Class I devices a post-market surveillance report. Article 86 assigns a PSUR to Classes IIa and IIb, as well as Class III, with class-specific update periods. The sample package should point to the applicable report and its supporting population.

Prove integrity and retrieval

A retention statement is configuration evidence. Article 10(8) requires technical documentation to remain available for at least 10 years after the last covered device reaches the market, rising to 15 years for implantable devices. The practical test retrieves an old controlled record and verifies that its links still resolve.

Ask records management to return one historical release packet and its risk file, followed by the attached field records. Record how long the query takes. Then copy one runtime event and alter the policy outcome before running the documented integrity check. Preserve the resulting failure output in the test file. The audit log chain-of-custody guide covers custody across joined systems.

My preference is blunt: an auditor should be able to break the package during a rehearsal. A missing join discovered internally becomes assigned remediation. The same gap found during notified-body review becomes the finding everyone remembers.

Separate request evidence from conformity evidence

For a manufacturer-controlled application that sends authenticated HTTP requests to an LLM, a request record can capture the supplied identity and workflow context. It can add destination and classification, plus the policy version. The action, response disposition, timestamp, and event identifier complete the row. Freeze all routed events for the review window before sampling. Include permits and redactions, followed by denials and processing failures.

That population excludes on-device inference and local models. Direct browser use plus opaque vendor-managed inference also sit outside it. Device qualification, MDR classification, clinical evaluation, and validation stay with qualified owners. CAPA decisions and retention policy do too, along with the conformity conclusion. AI governance for medical devices places those responsibilities in the wider lifecycle.

DeepInspect

DeepInspect sits on the authenticated HTTP path between a manufacturer-controlled application and an LLM endpoint. It evaluates application-supplied identity and approved workflow context, applies versioned request and response policy, and creates a signed, tamper-evident decision record for routed traffic.

Those records can supply a frozen event population and selected samples for one part of the technical file. DeepInspect leaves embedded inference, clinical validation, notified-body sampling, CAPA, retention decisions, and regulatory submissions with the manufacturer's designated owners. Book a technical deep dive at deepinspect.ai.

Frequently asked questions

Who should select the audit sample?

The notified body or independent reviewer should select it from a frozen device or event population, or approve a documented selection method. Manufacturer-picked successful cases introduce bias and provide weak completeness evidence.

Does MDCG 2025-6 prescribe an AI event sample size?

No universal count is given. The joint guidance says the governing MDR or IVDR technical-documentation sampling rules remain applicable. It does not create a universal count for runtime prompts, outputs, or model decisions.

What belongs in an assessor package for AI software?

Include the intended-purpose and classification anchor, approved configuration, risk and validation cross-references, clinical evaluation, post-market plan, selected field signals, change and CAPA dispositions, integrity results, and retrieval evidence.

How long must the technical documentation remain available?

Article 10(8) sets at least 10 years after the last covered device is placed on the market. The period is at least 15 years for implantable devices. Other applicable duties may require different records or longer handling.

Can request logs prove EU MDR conformity?

They can support traceability for authenticated HTTP traffic that crossed the governed route. The manufacturer and notified body rely on the complete technical file, including risk management, validation, clinical evidence, post-market surveillance, and quality-system records, for conformity work.