← Blog

Public Sector AI Audit Trails Answer the High-Impact Determination

Parminder Singh
Parminder Singh··5 min read
Summarize with AI

OMB Memorandum M-25-21, issued April 3, 2025, directs federal agencies to implement minimum risk management practices for high-impact AI, to establish a process for determining and documenting high-impact use cases, to measure and monitor ongoing performance and to centrally track those determinations. An audit trail is what makes those processes examinable. This article maps the memorandum duties to the per-request records an agency has to hold.

Industry Verticalsai-governanceai-compliancegovernmentauditnist-ai-rmf
Public Sector AI Audit Trails Answer the High-Impact Determination

OMB Memorandum M-25-21, signed by Director Russell T. Vought on April 3, 2025, directs agencies to implement minimum risk management practices for AI that could have significant impacts when deployed, which the memorandum calls "high-impact AI." It defines that category by effect: AI is high-impact when its output is a principal basis for decisions or actions that have a legal, material, binding, or significant effect on rights or safety. AI audit trail public sector work follows directly from that definition, because a determination about effect has to be evidenced against what the system actually did.

The memorandum also states that a high-impact determination "is possible whether there is or is not human oversight for the decision or action." A human in the loop does not remove the obligation, so the record has to show the AI step as well as the human one.

TL;DR

  • OMB M-25-21 was issued April 3, 2025 and directs agencies to apply minimum risk management practices to high-impact AI.
  • Chief AI Officers must establish a process for determining and documenting high-impact use cases and centrally track those determinations.
  • Agencies must establish processes to measure, monitor and evaluate ongoing performance and effectiveness of high-impact AI applications.
  • Where risk mitigation is not possible, agencies must cease use, which means the record has to support a discontinuation decision.

The determination needs an evidence base

The M-25-21 memorandum assigns the Chief AI Officer responsibility for "establishing a process for determining and documenting AI use cases as high-impact," for "establishing a process for an independent review of high-impact use cases before risk acceptance" and for "centrally tracking high-impact use cases and use case determinations."

A determination is a judgment about what an output is used for. Defending it later requires knowing which decisions the output actually fed. A use case documented as decision support and operating in practice as the sole basis for an eligibility outcome is a finding waiting to happen, and only request-level records distinguish the two.

The memorandum gave agencies 365 days from issuance to document implementation of the minimum practices and be prepared to report them to OMB, as part of periodic accountability reviews, the annual AI use case inventory or on request.

Monitoring duties need dated artifacts

The memorandum requires processes "to measure, monitor, and evaluate the ongoing performance and effectiveness" of high-impact AI applications, and it states that where high-impact AI is not performing at an appropriate level, agencies must have a plan to discontinue use until compliance is achieved. If proper risk mitigation is not possible, agencies must cease use of the AI.

Because discontinuation can interrupt a public service, the monitoring output behind it needs to be tied to specific releases and dates rather than a general assessment, with the population measured and the measurement method recorded.

The NIST AI Risk Management Framework organizes the same work under MEASURE and MANAGE, and most agency documentation already references it, which makes it a useful spine for the control narrative.

Use case inventories and request records answer different questions

Agencies other than the Department of Defense and the Intelligence Community must inventory AI use cases at least annually, submit the inventory to OMB and post a public version on the agency website.

An inventory entry is a description of a system. A request record is an observation of its operation. The inventory tells a reviewer that a benefits triage model exists and who owns it. The record tells a reviewer that on a given date a specific caseworker identity submitted a specific class of applicant data to that model and received a specific output under a named policy version.

Both are needed, and agencies tend to invest in the first because it has a submission deadline. AI governance in the public sector covers the program structure, and AI audit trail requirements by regulation covers field-level expectations across regimes.

Four fields for an agency record

The per-request record carries four things that an independent reviewer can test. The authenticated identity of the person or service that made the request, taken from the agency directory. The classification of the data in the request, which for agency work often means a privacy category under a system of records notice. The resolved destination model and its hosting environment. The policy version and the outcome.

Picture the independent review in practice. A reviewer asks for twenty sampled interactions from the quarter and wants to see, line by line, who submitted what and under which rule. An architecture diagram and a policy memorandum answer neither half of that request.

Independent review needs records the program did not write

The memorandum calls for independent review of high-impact use cases before risk acceptance. Independence is weakened when the only evidence available was produced and retained by the team whose work is under review.

Writing AI request records on a path the operating team cannot alter solves this cheaply. Signed audit logs for AI requests covers the write-path independence argument, and the same property supports an Inspector General review or a GAO engagement later.

Retention needs an explicit decision too, set against records schedules rather than against a default log retention setting chosen by a platform team.

Boundaries of the claim

An agency AI estate rarely sits in one place. Some inference runs in a cloud tenancy the agency controls, some inside a commercial SaaS product on an existing authorization and some on workstations. Only the HTTP paths the agency routes through its own enforcement point produce request records of this kind.

State that boundary in the control narrative with the excluded populations named and the alternative evidence for each one. A coverage percentage without named exclusions is the fastest way to turn a monitoring claim into a finding. AI gateway for government agencies covers the managed path in more detail.

DeepInspect

DeepInspect is a stateless proxy for authenticated HTTP traffic between agency users or agents and LLM endpoints. It evaluates application-supplied identity, request classification, approved destination and policy before forwarding, and every permit, redaction, reroute or block produces a signed per-decision record outside the calling application's write path.

For an agency, those records give an independent reviewer something to sample: which identity sent which class of data to which model under which rule, on a dated basis tied to a release. DeepInspect does not make the high-impact determination, assess model performance, cover vendor-native inference or replace an authorization to operate. Book a demo today.

Frequently asked questions

Which AI uses count as high-impact under M-25-21?

AI whose output is a principal basis for decisions or actions with a legal, material, binding or significant effect on rights or safety. The memorandum directs agencies to evaluate the specific output and its potential risks when assessing applicability, and it states that the determination is possible whether or not there is human oversight of the decision.

Does a human reviewer remove the obligation?

No. The memorandum states a high-impact determination is possible whether there is or is not human oversight for the decision or action. The practical consequence is that the record must capture the AI step and the human step separately, so a reviewer can see what the model produced and what the person did with it.

How does this relate to the annual use case inventory?

The inventory is a system-level submission to OMB with a public version posted on the agency website. Minimum risk management practices for high-impact use cases are operational duties evidenced by dated artifacts. Agencies are expected to be prepared to report implementation as part of periodic accountability reviews, the inventory or on request.

What evidence supports a decision to discontinue an AI use?

Monitoring results tied to specific releases and dates, the measured population and method, the threshold that was breached, the risk mitigation options considered and the named official who accepted or rejected continued use. The memorandum requires a plan to discontinue where high-impact AI is not performing at an appropriate level.