Accounting Firm AI Audit Trails Must Preserve the Work Behind the Workpaper
An accounting firm that uses an LLM to analyze a ledger, summarize a contract or draft an audit memo creates an evidence problem alongside the productivity gain. The engagement file may show the final workpaper while omitting the model, source data, policy and reviewer action that produced it. This article maps request-level AI records to PCAOB evidence expectations and the NIST AI Risk Management Framework without treating an AI log as audit evidence by itself.

On August 23, 2024, the SEC approved PCAOB amendments for audit procedures that use technology-assisted analysis of information in electronic form. The amendments apply to audits of financial statements for fiscal years beginning on or after December 15, 2025. They clarify the auditor's responsibilities for evaluating external information provided by the company, investigating items identified by a test of details and meeting every objective when one procedure serves more than one purpose.
That standard-setting action gives ai audit trail accounting firms work a practical target. If an auditor uses an LLM to classify journal entries or summarize a 70-page contract, the engagement file needs enough context to reconstruct the procedure. The polished memo alone hides the request that shaped it.
TL;DR
- An AI-assisted workpaper needs a record of the user, source data, destination model, policy decision and reviewer action.
- PCAOB technology-assisted analysis amendments keep responsibility with the auditor and sharpen evidence reliability duties.
- A request log supports reconstruction, but the engagement team still decides whether model output is competent evidence.
- Coverage claims must exclude desktop models and vendor-native AI traffic that never crosses the firm's HTTP gateway.
The final workpaper can hide the procedure
The SEC order approving the PCAOB technology-assisted analysis amendments describes principles-based changes intended to remain useful as audit technology changes. The order emphasizes controls over information technology, the reliability of external information provided by the company and enough disaggregation for information to remain relevant as audit evidence.
A workpaper drafted with an LLM can look complete while leaving those issues unresolved. A senior associate may paste a general ledger extract into a browser assistant, ask for unusual entries and carry five results into a testing sheet. The sheet records the five entries. It may omit the original prompt, excluded columns, model endpoint and twenty other entries the model returned. On review day, the manager sees a clean spreadsheet with green tick marks and no view of the selection mechanism.
I would reject a file that records the answer while losing the procedure that produced it.
Five fields connect an AI request to the engagement file
A useful per-request record starts with the authenticated employee or service identity. Shared application credentials reduce this to a generic account, which prevents the engagement partner from establishing who initiated the analysis. The second field is a classification of the submitted material: public filing, client confidential data or personally identifiable information, for example.
The record then needs the resolved model and hosting environment, because a firm-approved Azure OpenAI deployment and a personal web account create different custody questions. A fourth field identifies the policy version and its decision, including any redaction or block. The last field binds the interaction to an engagement, workpaper reference or approved use case supplied by the calling application.
Together, these fields create provenance while leaving professional judgment with the engagement team. The AI audit trail requirements by regulation article explains the same separation between operational events and management-level records across regulated sectors.
Evidence reliability stays with the auditor
The PCAOB amendments clarify that an auditor evaluating information received from an external source and provided by the company must assess its relevance and reliability. An LLM adds another processing step. It can reorder, omit or infer content before the auditor sees the output, so the team needs a documented method for checking the result against the source population.
For a contract review, that method might require the reviewer to compare every cited clause with the signed agreement and record exceptions in the workpaper. For a journal-entry analysis, it might require reconciliation of the submitted row count and control totals, preservation of query parameters and follow-up on every item that meets the test criteria. The AI interaction record proves which request reached which model. It cannot prove that the response was right.
A fluent paragraph on the screen is presentation. Evidence requires the source and the checks behind it. AI governance for accounting firms covers ownership and approval at the use-case level; the request record tests whether staff followed that approval in practice.
NIST turns governance into dated records
The NIST AI Risk Management Framework 1.0 places inventory, defined roles and ongoing monitoring inside the GOVERN function. Its subcategory 1.5 calls for planned monitoring and periodic review of the risk-management process; subcategory 1.6 addresses AI-system inventory according to organizational risk priorities. Documented responsibility and communication lines for mapping, measuring and managing AI risk appear in subcategory 2.1.
An accounting firm can map those outcomes to engagement operations. Start with an approved-use inventory that names the procedure and owner, then compare it with actual request records each quarter. Assign exceptions to a partner or risk leader. If tax staff begin sending client trial balances to a model approved only for public-company research, the interaction population exposes the drift.
The useful report is concrete: policy version 14 blocked a client-data request at 10:42 a.m., the engagement code was ACME-2026 and the risk owner closed the exception on October 4. A slide saying "AI use monitored" carries none of that weight.
Retention must follow the engagement obligation
Accounting firms already retain engagement records under professional, legal and contractual rules. AI gateway logs often arrive with a platform default chosen for operational troubleshooting. Those two clocks need an explicit mapping. Deleting request records after 30 days can leave a seven-year engagement file with no trace of the AI procedure behind a sampled workpaper.
The better design stores a stable reference in the workpaper and preserves the linked request record under the applicable engagement schedule. The record should also survive outside the application that called the model. Signed audit logs for AI requests explains how signatures and a separate write path expose later modification.
Retention requires restraint as well. Prompts can contain full ledgers, tax identifiers and contract language. The audit record can store classifications, hashes, decision metadata and approved references while placing sensitive payloads in controlled evidence storage. The firm's records policy should name the system of record, access role and disposal trigger for each field.
The HTTP boundary belongs in the control description
An HTTP enforcement point can record browser and application calls routed through it. It misses a local model running on an auditor's laptop, an embedded feature that performs inference inside an audit platform and any personal device outside firm management. Vendor-native processing may produce its own logs, but those records sit under the vendor's schema and custody.
A defensible control description lists those exclusions. For example, the firm's population may cover managed-browser traffic to approved model endpoints and an internal audit application, while a desktop document assistant remains under a separate application control. AI vendor risk for accounting firms addresses the evidence requested from that second group.
A broad claim such as "all AI use is logged" invites an inspector to find the one laptop with a model running beside a folder of client PDFs. Name the covered routes, test them each quarter and keep the test result with the policy version.
DeepInspect
DeepInspect is a stateless proxy for authenticated HTTP traffic between firm users or agents and LLM endpoints. It evaluates application-supplied identity, request classification, approved destination and policy before forwarding. Each permit, redaction, reroute or block produces a signed per-decision record outside the calling application's write path.
For an accounting firm, those records can tie an engagement identity and data class to the model and policy used for a specific procedure. DeepInspect does not judge audit evidence, approve a workpaper, cover local inference or replace the firm's quality-control system. Book a demo today.
Frequently asked questions
- Does an AI request log become audit evidence?
The log is evidence of the interaction: which identity submitted a classified payload, which destination received it, what policy applied and what decision occurred. The engagement team still evaluates the source information, model output and procedure under the applicable auditing standard. For a December 2026 journal-entry test, the workpaper should tie the logged request to the population, selection criteria, follow-up and reviewer sign-off. A cryptographic signature protects the record's integrity; it does not make the model's analysis competent or sufficient.
- Should the firm store complete prompts and responses?
The answer depends on the engagement requirement and the data involved. Full payload storage may expose client ledgers, taxpayer identifiers or privileged communications to a larger log-review population. A safer design records identity, classification, model, timestamp, policy outcome and a hash or controlled evidence reference. The engagement repository can retain the payload under tighter access. The October 2026 records schedule should state which system holds each element and how long it remains available.
- What should an engagement partner review?
The partner needs the approved use case, the population of AI interactions tied to the engagement and the exceptions. A sample should connect each request to a workpaper, named reviewer and policy version. For technology-assisted analysis under the PCAOB amendments effective for fiscal years beginning on or after December 15, 2025, the file should also show how the team evaluated source reliability and investigated items that required follow-up.
- How should a firm handle AI built into an audit platform?
Treat vendor-native inference as a separate evidence population. Ask the vendor for user identity, model and version, source lineage, retention controls, administrative access and export capability. Then document which fields remain unavailable. An external HTTP gateway covers calls that the firm routes through it; inference wholly inside the vendor platform needs vendor records and application controls. The control narrative should keep those populations separate rather than present one coverage percentage.