Traditional software fails loudly. It throws an error, the status code changes, and a dashboard turns red.

Key Takeaways
  • Instrument model and especially tool layers: capture calls, arguments, responses, errors, retries, durations, and version metadata to reconstruct failures.
  • Treat a trace as the primary unit: record LLM calls, tool invocations, and retrievals as child spans to reconstruct the agent's reasoning.
  • Pair OpenTelemetry with a production evaluation layer and a preapproved content policy: sample scoring, redaction rules, retention, and environment permissions.

Agents fail successfully. They complete the task, return a 200, and do the wrong thing — calling the wrong tool, retrieving the wrong document, or confidently producing an answer built on a failed lookup nobody noticed.

That single difference explains why teams instrument agents with their existing monitoring stack, see healthy latency percentiles and error rates within a week, and still cannot answer why the system did what it did.

Here is what agent observability actually requires.

Why Your Existing APM Falls Short

Four properties of agent systems break the assumptions monitoring tools were built on.

Non-determinism. The same prompt produces different outputs. You cannot reproduce an issue without having captured the exact input, model parameters, and settings at the time of the call. If you did not record them, the incident is unreconstructable.

Multi-step opacity. A single user request may involve a dozen model calls, tool invocations, and retrieval steps. Aggregate latency tells you the request was slow. It does not tell you which of the twelve steps was slow, or why the agent chose that path.

Token-based economics. Cost and latency correlate with token counts rather than request counts. A dashboard counting requests is measuring the wrong unit entirely.

Content sensitivity. Prompts routinely contain personal, confidential, or regulated data. Shipping them to an observability backend without sanitisation converts a monitoring decision into a compliance problem.

The Failures Are Not in the Model

The most useful correction available, and it reorients where you spend instrumentation effort.

Practitioners working on production agent systems consistently report that the model is rarely the root cause. The agent selects the right tool and calls it correctly — and the tool fails. Authentication expires, a rate limit trips, a request is malformed, an API returns an unexpected shape.

See also  AI In Property Valuation: Accurate & Fast Real Estate Pricing

As one observability team put it, “the model said something weird” is the least useful line available in an incident postmortem. It is usually also wrong.

The practical consequence: instrument the tool layer as carefully as the model layer. You need to see not just that the agent called a tool and received an error, but what it sent, what came back, and what it did next.

What to Instrument

Five layers, at minimum.

1. LLM calls. Model, parameters, prompt, completion, token counts, finish reason, latency, cost.

2. Tool invocations. Which tool, what arguments, what response, success or failure, retries, duration. This is where most failures live and where instrumentation is most often thin.

3. Retrieval steps. What query, which sources, what was returned, what was actually used. A retrieval that returns nothing and an agent that proceeds anyway is a silent failure with no error attached.

4. Application metadata. User, session, tenant, agent version, prompt version, environment. Without version metadata you cannot correlate a behaviour change to the change that caused it.

5. Decision points. Where the agent chose between options — which branch, and on what basis.

That last one requires manual work. Auto-instrumentation captures the standard layers well; it does not know where your agent’s meaningful decisions are. Treat auto-instrumentation as the floor, then add manual spans at your own decision points.

The Trace Is the Unit

The structural shift: for agents, the useful primitive is a trace, not a log line or a metric.

Each LLM call, tool invocation, and retrieval step becomes a child span within a single trace representing one agent run. Read in order, that trace reconstructs the reasoning chain — what the agent knew, what it tried, what came back, and what it did next.

Metrics tell you something is wrong across a population. A trace tells you what happened in one case. For non-deterministic systems, the second is where debugging actually occurs.

The Standards Are Emerging, Not Settled

Worth being precise about, because tooling decisions depend on it.

OpenTelemetry GenAI semantic conventions are the community standard for vendor-neutral instrumentation, defining consistent attributes such as gen_ai.system, gen_ai.request.model, gen_ai.usage.input_tokens, and gen_ai.usage.output_tokens. Major observability vendors have added native support, and auto-instrumentation packages exist for common providers and frameworks.

See also  10 AI Tools for HR Analytics That Turn People Data Into Action

But the conventions are not finished. In June 2026 the project moved GenAI, provider-specific, and MCP conventions into a dedicated repository so they could version independently. As of late August 2026 that repository marked the GenAI conventions as being in development, without an official release.

OpenInference is an alternative set of conventions that some practitioners prefer for production agent workloads, on the grounds that it captures richer detail — full prompt and completion content, cost, and model parameter metadata — where the community standard is still maturing.

The practical advice: adopt OpenTelemetry as the transport and pipeline regardless, since it lets you redact, enrich, and route telemetry before it leaves your network. Then pin your convention version and treat it as a contract, because silent schema drift will break dashboards without anyone noticing.

What the Standards Do Not Cover

The gap that catches teams out.

The GenAI conventions standardise how you capture model attributes, token usage, and latency. They do not cover output evaluation, safety scoring, or content quality assessment. Conventions for embedding evaluation results directly into spans remain in early discussion.

So a complete setup has two parts: OpenTelemetry as the data plane, and a separate evaluation layer that scores outputs for correctness, faithfulness, policy compliance, and whatever else matters in your domain.

Without the second part you have excellent visibility into a system doing the wrong thing efficiently.

Evaluation in Production

Three practices that make this workable rather than theoretical.

Sample, do not score everything. Running automated judging on every production trace is expensive and unnecessary. Scoring 10 to 20% of traffic is a commonly cited balance between quality coverage and evaluation cost.

Build golden datasets from production. Promote interesting real traces — failures, edge cases, high-value successes — into curated, version-controlled test sets. Those become the regression suite you run before every prompt or model change, and they are far more representative than synthetic examples.

Test changes against them. Prompt changes, model upgrades, and agent strategy changes should be evaluated against the same dataset with the same scoring, so you can attribute a performance shift to a specific change.

Handle Content Before You Turn It On

A compliance decision disguised as a configuration flag.

See also  10 AI Tools for Supply Chain Management That Deliver Smarter Operations

Capturing full prompts and completions makes debugging dramatically easier and puts potentially sensitive content into your observability pipeline. Decide the policy first:

  • Which environments may capture full content, and which capture metadata only
  • What is redacted or truncated at the instrumentation layer, before export
  • Where telemetry is permitted to be sent, and under whose control
  • How long it is retained
  • Who can retrieve it

Sanitise in the instrumentation wrapper rather than in the backend. Data that never leaves your boundary is easier to defend than data you filtered after arrival.

The Metrics Worth Tracking

  • Task completion rate — did the agent finish, and correctly? Separate the two.
  • Tool call success rate, per tool. This surfaces the failures error dashboards hide.
  • Steps per task, watched for drift. Rising step counts usually mean degrading efficiency or looping.
  • Token consumption per task, which is your real cost unit.
  • Retry rate, a leading indicator of both cost and reliability problems.
  • Latency by span type, not just end to end.
  • Evaluation scores over time, on your sampled traffic.
  • Human intervention rate — how often someone has to step in.

That last one is the closest thing to a single health metric. If it is rising, something upstream is degrading.

Common Mistakes

  • Monitoring the model, not the tools. The failures are mostly downstream.
  • Relying on error rates. Agents fail with successful status codes.
  • Auto-instrumentation only, with no manual spans at decision points.
  • No version metadata, making it impossible to correlate behaviour changes to releases.
  • Capturing content without a policy, which creates an exposure you did not intend.
  • Treating observability as a launch task. Agent behaviour drifts as models, prompts, and tools change around it.

Final Thoughts

Agent observability is not application monitoring with extra fields. It is the discipline of reconstructing why a non-deterministic, multi-step system made the choices it did.

Instrument the tool layer as seriously as the model layer, make the trace your primary unit, pair OpenTelemetry with a real evaluation layer, and decide your content policy before you turn on prompt capture.

The teams that get this right early are the ones that ship agents with confidence. The ones that skip it tend to stay in the demo phase, because they cannot prove the thing works.

How useful was this post?

Rated 0 / 5. Vote Count: 0

Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?