AI Agent Observability: Traces, Decisions, and Debuggable Behavior
AI agent observability needs a trace for each model call, tool handoff, retrieval result, and policy decision. This guide defines the signals, OpenTelemetry fields, correlation IDs, and retention split that make autonomous behavior reviewable.

An AI agent that runs a multistep workflow and returns a result to the calling application produces one trace in most application observability stacks: the entry-point request and its final response. Everything the agent did between those points, including retrieval and policy or token records, collapses into an opaque black box. When a customer support agent gives a refund to the wrong account, the incident review team has one signal to work with: the fact that a refund happened. The signals that would answer "why" are missing from the telemetry pipeline. I want to walk through the observability signals AI agents have to emit, the OpenTelemetry semantic conventions taking shape, and where the AI request boundary sits in the pipeline.
TL;DR
- AI agent observability needs a trace that joins model requests and policy-decision evidence to the agent and tool-execution record.
DeepInspect contributes evidence only for routed HTTP model requests. Tool calls, retrieval systems, and local execution need instrumentation at their own boundaries.
Application observability treats the agent as one call. Agent observability treats each internal step as a call.
The signals that matter
Six signal categories separate a debuggable agent from a black box.
Tool call traces. Every tool the agent selects and every tool call it issues has to land in the trace with its name, arguments, and result. When a coding agent runs git commit with the wrong scope or a support agent calls the issue_refund tool with a customer ID it derived incorrectly, the trace shows the argument the agent chose, not just the effect.
Delegation traces. In multi-agent workflows, the parent agent hands sub-tasks to child agents. The trace has to connect the parent's decision to delegate with the child's execution, so the incident review can follow the delegation chain.
Policy decisions. Every policy evaluation at the AI request boundary produces a signal: the policy version, identity claim, and allow-or-deny outcome. The policy-as-code piece covers the artifact side; observability covers the runtime side.
Retrieval hits. For agents backed by RAG, the retrieval step selects context that shapes the model's output. The trace has to include the retrieval query and the selected documents with their similarity scores. When the agent hallucinates from retrieved content, the retrieval trace shows what the model actually saw.
Token consumption per step. Cost attribution runs on per-step token counts, not per-request totals. An agent that runs 40 tool calls in a session consumes 40 sets of tokens; the aggregate hides which step is expensive.
Content classification tags. When an LLM request or response payload triggers a classifier, such as PII, PHI, prompt injection heuristics, or jailbreak patterns, record that classification in the trace so incident review can filter high-risk sessions.
OpenTelemetry semantic conventions for GenAI
The OpenTelemetry GenAI semantic conventions define span attributes for LLM calls, including gen_ai.system, gen_ai.request.model, gen_ai.usage.input_tokens, and gen_ai.usage.output_tokens. W3C's Trace Context specification supplies the traceparent header that lets the model-call span, application span, and policy-decision record refer to the same request.
The conventions matter because they make traces portable across observability vendors. A trace emitted with OTel GenAI fields flows into Datadog, Grafana, Honeycomb, or a custom SIEM without vendor-specific transformation. The AI request logging format covers the mapping between the AI audit log format and the OTel semantic conventions.
The gap in the current conventions is on the policy and identity side. OTel GenAI conventions cover what the model did. They do not, yet, cover which identity the request was bound to or which policy decision applied. That gap is where AI-security-specific instrumentation extends the standard set.
The pipeline
The production pipeline for AI agent observability has three stages.
The first stage emits records. Each component in the agent stack, including the framework runtime, tool implementations, retrieval layer, policy engine, and AI gateway, emits OTel spans with GenAI fields. Framework instrumentation usually covers model-call spans. The policy engine and gateway add identity and policy fields.
The second stage is collection and enrichment. An OTel collector receives spans, adds environment metadata such as deployment tag, region, and tenant, then routes records to the selected telemetry store. Enrichment at this stage lets operations query by tenant or deployment version without every emitter carrying its own context.
The final stage analyzes the records. The telemetry store answers operational queries for per-tenant token spend and error or denial rates. Security review uses the same data for classifier triggers and delegation chains crossing trust boundaries. The AI gateway observability guide covers the operational side.
Correlation and retention
The trace ID has to cross the application and routed LLM request without becoming an untrusted label supplied by the model. Generate it at the application edge, propagate it with W3C Trace Context, and store the identity binding and policy version beside it at the enforcement point. Tool servers and local execution need their own instrumentation and authorization controls; a gateway cannot observe calls that never traverse its HTTP model-request path.
Keep high-cardinality traces in the operational store for a short diagnostic window, then retain the signed policy-decision record under the evidence-retention rule. A viewer should be able to start with a denied request on a gray Monday morning and locate the exact policy version without retaining every prompt indefinitely.
Where the AI request boundary fits
The agent framework alone cannot produce the identity and policy signals. The framework knows the model call happened; it does not know whether the calling identity was authorized, which policy version applied, or which classification the payload carried. Those signals originate at the AI request boundary where identity binding and policy evaluation happen.
The gateway can emit its own spans directly to the telemetry pipeline and join them to the agent framework by trace ID. An SDK integration can also attach the gateway span during the framework call. Team ownership of the framework and gateway determines the better integration route.
DeepInspect
This is exactly what DeepInspect does for routed HTTP traffic between authenticated users or agents and an LLM. DeepInspect sits at that AI request boundary as an external enforcement layer. Each routed model request can emit OpenTelemetry spans with GenAI fields plus identity claims, policy version, and content classification. Those spans flow into the customer's OTel collector and selected telemetry store.
Trace IDs link the agent framework, gateway, and LLM-provider response for the model-request segment. Tool calls require instrumentation at their own execution boundary. The durable, tamper-evident audit log references the same trace ID for cross-referencing.
Book a technical deep dive at deepinspect.ai.
Frequently asked questions
- How is AI agent observability different from LLM observability?
LLM observability covers the model call: prompt, response, token counts, latency. Agent observability covers everything around the model call: tool selection, tool arguments, sub-agent delegation, retrieval, policy evaluation, identity binding. A production agent produces dozens of spans per user request, while one LLM call produces one span.
- Does OpenTelemetry cover this?
Partially. The OTel GenAI semantic conventions cover LLM calls and are extending to agents. Identity binding and policy decisions are not yet part of the standard set. Production deployments extend the OTel attribute set with organization-specific attributes for the gaps.
- How much telemetry data does an agent produce?
An agent that runs 40 tool calls in a session produces 40 tool call spans plus surrounding spans for retrieval, policy, and delegation. A high-throughput agent deployment can produce hundreds of gigabytes of trace data per day. Sampling strategies (head sampling for cost-sensitive traces, tail sampling for error and high-latency traces) apply the same way they apply for application traces.
- Where should the observability data live?
Traces belong in the telemetry platform the operations team already runs, such as Datadog, Grafana Tempo, Honeycomb, or self-hosted Jaeger. The tamper-evident audit log belongs in an append-only store separate from that platform, because the audit record has to survive a telemetry outage or compromise. The AI audit log immutability guide covers the storage-layer contract.
- How do we link observability data with audit logs?
Trace ID propagation. Every span in the observability pipeline and every record in the audit log carries the same trace ID for a given user request. Incident review starts in either place (a suspicious span in the trace, an unusual policy decision in the audit log) and pivots to the other by trace ID.
- What is the latency cost of emitting these signals?
Sub-millisecond per span with a local OTel collector and async export. The dominant latency in the AI request path is the model call itself (hundreds of milliseconds to seconds). Observability adds under 1% to that total. The ai gateway latency piece covers the measurement methodology.