ISO/IEC 5338 AI Audit Evidence Across the System Life Cycle
ISO/IEC 5338:2023 defines AI system life cycle processes covering acquisition, organizational support, technical management, engineering, operation, maintenance, and disposal. This guide shows how to build an evidence chain across those processes, connect design decisions to production AI requests, and keep the distinction between process conformance and ISO/IEC 42001 management-system certification clear.

ISO/IEC 5338:2023 starts with a process model. It extends established system and software life cycle processes for AI systems based on machine learning and heuristic methods, then groups the work into agreement, organizational project-enabling, technical management, and technical processes. The standard reaches past model development into acquisition, operation, maintenance, continuous validation, and disposal.
That scope changes the audit question. I want to see a traceable chain connecting an approved use case, requirements, data and model decisions, production controls, observed requests, change records, and retirement actions. A lifecycle binder with no production request sample is theater with page numbers.
Establish the standard and certification boundary
ISO/IEC 5338:2023 defines processes and associated concepts for the AI system life cycle. Its scope covers the definition, control, management, execution, and improvement of an AI system, including projects that develop or acquire one. The standard contains a conformance concept, but it is a life cycle process standard rather than the management-system certification target used for an accredited AI management system audit.
ISO/IEC 42001:2023 fills that certification role by specifying requirements for an AI management system. A company can use 5338 to design and operate its lifecycle processes, then use the resulting records as evidence inside a 42001 program. The certificate, when sought, attaches to the management system and its declared scope.
Preserve the exact 5338 edition, the chosen conformance approach, the systems in scope, and the approval date. Put those four fields on the mapping cover sheet. An assessor should never have to infer which edition a control matrix used.
Build an evidence chain around the four process groups
The agreement processes cover acquisition and supply. Evidence should identify the acquired model or AI service, supplier, approved purpose, service terms, required data handling, acceptance criteria, and exit obligations. For an external LLM, include the permitted endpoint and the contract version reviewed by legal or procurement.
Organizational project-enabling processes cover life cycle model management, infrastructure, portfolio, people, quality, and knowledge. Keep the adopted lifecycle model, role assignments, competence records, environment inventory, quality plan, and knowledge-transfer artifacts. A RACI chart alone leaves the operating mechanism unclear. Pair each named owner with an approval or review they actually performed.
Technical management processes include planning, assessment and control, decision management, risk, configuration, information, measurement, and quality assurance. The evidence chain needs dated baselines, decision records, risk treatments, configuration history, measurement definitions, review results, and corrective actions.
Technical processes carry the chain into requirements, architecture, design, data engineering, implementation, verification, transition, validation, continuous validation, operation, maintenance, and disposal. This is where production telemetry becomes indispensable.
Connect requirements to production AI traffic
Take one requirement such as “customer account numbers may reach only the approved private model route after redaction.” Its evidence chain should contain a requirement ID, architecture decision, data classification, policy rule, test case, deployment approval, and production samples showing permit, redact, and deny outcomes.
A request-level record should include the request ID, timestamp, verified caller identity, delegated user where applicable, system and route IDs, destination model, data classification, policy version, decision, and enforcement action. Add a tamper-evident signature or equivalent integrity mechanism. Keep the evidence write path outside the application that generated the model call.
The physical test is simple. Place the approved requirement on the left side of an assessor’s desk and one signed JSON decision record on the right. The control owner should be able to explain every link between them without opening six dashboards and guessing which policy was live.
Application logs can contribute context, but application custody creates the self-attestation problem. A crash after the model responds, a disabled logger, or an administrator with delete access can break the chain. Independent records make the production control testable.
Preserve change and continuous-validation evidence
ISO/IEC 5338 names configuration management, measurement, verification, validation, continuous validation, operation, and maintenance as separate lifecycle processes. Treat a model swap, policy change, prompt-template release, new data category, endpoint change, and classifier update as controlled changes with distinct evidence.
For each change, retain the request, owner, affected system IDs, risk review, approval, pre-deployment result, release identifier, effective time, rollback condition, and post-deployment observation. The first production record after the change should identify the new model or policy version. That single correlation catches a common gap: the change ticket says version 14 shipped, while runtime records still show version 13.
Continuous validation needs defined measures and thresholds. The AI development team may own model quality, drift, bias, and task performance. The request boundary owns observable traffic properties such as unauthorized destinations, sensitive-data decisions, policy failures, and unexplained identity gaps. Keep these scopes separate in the evidence plan.
When a threshold fails, preserve the alert, triage decision, containment action, retest, and closure approval. A green dashboard screenshot proves very little unless the team can reproduce the underlying query and time range.
Cover suppliers, maintenance, and disposal
An acquired AI service remains inside the 5338 lifecycle scope. Keep supplier evaluation, model or service version notices, contract changes, endpoint approvals, incident notifications, and acceptance results under the same stable system ID used for internal evidence. Procurement folders and runtime logs should point to the same object.
Maintenance records need the defect or change reason, affected configuration items, validation results, release decision, and observed production outcome. If the provider silently changes a model behind a stable API name, record the date detected, the evidence used to identify the change, and the reassessment decision.
Disposal deserves its own evidence package. Record the retirement approval, traffic cutoff, credential revocation, route removal, retained evidence, deletion actions, supplier termination, and final verification. Capture a denied request against the retired route after cutoff. That concrete test proves the old endpoint stopped receiving production traffic.
A dusty architecture diagram with a red X through the model box is not disposal evidence. The final package should prove that access ended, data handling obligations were executed, and remaining records follow the approved retention schedule.
Assemble a reproducible audit package
Build the package around traceability rather than document volume. Include the scoped inventory, lifecycle model, supplier register, requirement baseline, architecture, risk and decision records, data-engineering artifacts, test plans, validation results, configuration history, runtime decision samples, maintenance log, exception register, and disposal evidence.
For every sample, record the population, date range, selection method, query, record identifiers, and hashes. Select at least one permitted request, one redacted request, one denied request, one exception, and one changed configuration. The assessor should be able to run the same query and retrieve the same evidence.
Use a manifest with six columns: 5338 process, control objective, owner, implementation, evidence location, and last test. Add a seventh column for gaps. Blank coverage is more credible than a vague claim that one control covers acquisition, data engineering, continuous validation, and disposal.
The useful outcome is a maintained chain between lifecycle work and observed operation. ISO 5338 gives the process structure. Your evidence system proves that the structure operated on a specific AI system at a specific time.
DeepInspect
DeepInspect covers the HTTP AI request boundary inside that larger lifecycle. It sits between authenticated users or agents and LLM endpoints, evaluates identity, role, prompt classification, route, model authorization, and policy, then records the decision before forwarding permitted traffic.
Each permit, redact, and deny outcome produces a tamper-evident record with the policy version and enforcement result. Those records support 5338 operation, configuration, measurement, quality assurance, continuous validation, maintenance, and incident reconstruction. DeepInspect leaves model training, data-pipeline quality, supplier contracting, human-resource processes, and disposal governance to the teams that own them.
Book a technical deep dive at deepinspect.ai.