IRS Publication 1075 AI Audit Evidence for Model Requests and Decisions
IRS Publication 1075 AI audit evidence starts with a complete population of authenticated model requests, then binds reviewer-selected samples to identity, destination, policy, decision, response handling, integrity, and retention. This guide separates evidence an HTTP policy gateway can produce from programme, legal, endpoint, and operational records that stay outside that boundary.

The evidence folder contains a red FTI cover sheet, a UTC export, and a denied prompt showing a nine-digit taxpayer identifier replaced before transmission. That single screen determines if an agency or contractor, including a consolidated data center, receiving or processing federal tax information has a usable control record. The IRS Publication 1075 sets the programme obligation, while IRS Safeguards Program supplies the control or assessment method. A irs publication 1075 AI audit evidence should connect those requirements to the authenticated HTTP request before the payload reaches an LLM.
I would rather show an IRS reviewer a denied test with its rule version than another policy page saying employees should be careful.
TL;DR
- Scope every AI route that can touch regulated data or a financially relevant workflow, including the calling application and supplied identity, plus the model destination and bypass paths.
- Test authorization and information-flow policy. Test event content and record integrity, along with historical retrieval and exception handling, using real requests rather than screenshots of configuration pages.
- Keep the boundary explicit. A gateway contributes to authenticated HTTP model traffic. Legal authority to receive FTI, disclosure accounting and employee awareness remain with their existing owners, as do physical storage, incident notification and the safeguard review itself.
- Retain the policy version and decision beside the event so a reviewer can reproduce what happened at that point in time.
Define the population before selecting samples
The population should include every routed model request in the assessment window. This includes permitted calls and redactions, as well as denials, missing-identity failures, policy errors and blocked responses. Export a manifest with the stable event identifier and UTC timestamp. Include the originating principal, calling application and model destination, followed by the classification, policy version, action and record location.
Freeze the manifest before the reviewer selects rows. Owner-picked screenshots demonstrate the product interface. A frozen population gives the reviewer a basis for completeness and selection. IRS Publication 1075 supplies the governing programme context, and IRS Safeguards Program supplies the control or testing language that the package should address.
Bind each sample to the authorization decision
A selected event needs the upstream identity assertion and validation result, plus the requested model route and data classification. It also needs the policy rule and version, the enforcement outcome and the response disposition. Join upstream authentication to the outbound call with a stable correlation identifier. A provider record showing one shared credential leaves the human or agent originator unresolved.
Use three samples with different outcomes and explain the expected result before inspection. The AI audit trail requirements guide gives the cross-regulation record fields, while NIST SP 800-53 Revision 5 anchors the framework-specific control method.
Show custody and integrity rather than asserting them
Name the writer and storage destination, the roles with write or delete authority, and the integrity mechanism. Then copy one staged event and change its decision field. Run the documented integrity check and retain the failure output. Attempt a write as an application administrator and retain the rejection.
A diagram with arrows helps here: application to policy point and policy point to model, with a separate arrow to the protected record store. The application should lack permission to rewrite that final box. The audit log chain of custody guide covers correlation and transfer between systems.
Test historical retrieval with the oldest useful record
A retention setting is configuration evidence. Retrieval evidence comes from restoring or querying an old event, validating its integrity, and joining it to the policy and identity records that existed at the time. Record the query and elapsed time, the archive step and returned count, plus the integrity result and reviewer.
The assessment team should repeat the query independently. If a storage tier requires restoration, that procedure belongs in the package and in the audit calendar. A screenshot reading "retained" proves less than a returned row with a matching integrity value.
Connect exceptions to human review
An audit package needs more than clean events. Include denied requests and exception approvals, along with emergency changes, rule failures and investigation dispositions. For one exception, trace the request and ticket, the approver, the start and end time, the policy version and resulting events, and the closure review.
This is where the agency safeguard lead and tax information owner divide accountability with the information security officer, IAM lead and AI platform owner. The request-layer evidence can show what policy did. Human approvals and legal conclusions come from separate systems and named owners, as do control deficiency grading and remediation decisions.
State the coverage boundary in the package
The governed population covers authenticated HTTP traffic routed through the control point. Direct consumer browser use and local inference remain outside it, alongside unmanaged endpoints, calls made with stolen provider credentials and inference hidden inside a software vendor. List those paths and the controls assigned to them.
Legal authority to receive fti, disclosure accounting and employee awareness require their own records. The same applies to physical storage, incident notification and the safeguard review itself. My preference is to place the boundary statement beside the population definition, because a reviewer should see the denominator and its exclusions before reading a perfect sample.
DeepInspect
DeepInspect intercepts authenticated HTTP AI traffic on its way to and from an LLM. It evaluates application-supplied identity and policy context, applies a permit or deny action with redaction where required, inspects the response, and commits a per-decision record to support complete populations and reviewer-selected samples.
Identity proofing and endpoint security remain with their established owners, alongside accounting or legal judgment and incident submission. Legal authority to receive FTI, disclosure accounting and employee awareness also remain with those owners, as do physical storage, incident notification and the safeguard review itself. Book a technical deep dive at deepinspect.ai.
Frequently asked questions
- Which events belong in the audit population?
Every in-scope routed request belongs. This includes permits and redactions, along with denials, timeouts, policy errors and missing-identity failures. Excluding failures changes the denominator and weakens completeness testing.
- Can prompt text be replaced with a fingerprint?
A fingerprint can support integrity and correlation while reducing sensitive content in the audit store. The package still needs a controlled route to the source record when a reviewer must inspect content, plus documentation describing what the fingerprint covers.
- Who should choose the audit sample?
The reviewer should select rows from a frozen population or approve a repeatable selection method. Control owners can explain records and provide context, but a hand-picked set of successful calls creates selection bias.
- What can the gateway evidence prove?
It can prove the supplied identity and context, plus the destination and policy version. It can also prove the classification and decision, response handling and record integrity for a routed request. Legal authority to receive fti, disclosure accounting and employee awareness remain separate conclusions backed by separate evidence. The same applies to physical storage, incident notification and the safeguard review itself.