edidiong umana · writing
home
Engineering notes4 min read

Traces: what your AI system actually did

A working guide to tracing AI features in production: spans and OpenTelemetry's GenAI names, what to keep out of your logs, the few numbers worth an alert, and how a production failure becomes tomorrow's test case.

Your dashboard says the average reply time is under two seconds. Users say the bot sometimes never answers. Both can be true at once. If 90 requests take 0.8 seconds and 10 take 12 seconds, the average is 1.92 seconds, and one user in ten waits twelve seconds for a reply.

An average can't show you that, and a log line can't tell you why. Once real users arrive, your test set stops being the whole story. You need a record of what the system did for each of them: a trace.

A trace is a tree of spans

A trace is the complete record of one request, from the user's first message to the final reply, across every model call and tool. OpenTelemetry, the open standard most tracing tools share, models it as a tree of spans. Each span is one timed operation with a name, some attributes and a status.

Here is an illustrative trace of a parcel-tracking support bot, with made-up values:

invoke_agent tiwa                      6.4 s   ERROR
│   gen_ai.operation.name = invoke_agent
│   gen_ai.provider.name  = ollama
│   app.feature           = tracking
│   app.user_hash         = u_7c1e9a          (pseudonymised, no phone number)
│
├── chat qwen3:4b-instruct             1.1 s   OK
│       gen_ai.request.model       = qwen3:4b-instruct
│       gen_ai.usage.input_tokens  = 812
│       gen_ai.usage.output_tokens = 46
│
└── execute_tool track_parcel          5.0 s   ERROR
        gen_ai.tool.name = track_parcel
        error.type       = timeout

Read it top to bottom. The model call was quick and cheap. The courier lookup took five seconds and timed out, and no second model call followed, so the customer never got a reply. The model wasn't the problem: the tool was. A fix might be a time limit on the lookup with an honest fallback message, plus a new test case for a lookup that times out.

Use the standard names, and check them

OpenTelemetry has semantic conventions for generative AI that give these attributes standard names, so any compatible tool can read your traces. The main ones:

What to recordStandard name
Kind of operationgen_ai.operation.name: chat, execute_tool, invoke_agent
Who served it (required)gen_ai.provider.name
Which modelgen_ai.request.model, gen_ai.response.model
Tokens usedgen_ai.usage.input_tokens, gen_ai.usage.output_tokens
Why the reply endedgen_ai.response.finish_reasons
Which toolgen_ai.tool.name
What went wrongerror.type

Two cautions. First, these conventions are still marked Development, and they move. gen_ai.system was renamed gen_ai.provider.name in 2025, and in June 2026 the GenAI conventions moved out of the main semantic conventions repository into their own. Check the current names before you build dashboards on them.

Second, cost isn't one of them. The conventions record tokens, not money. Work out cost yourself from token counts and your own price table, per feature and per user. Put your own attributes in your own namespace, such as app.feature or app.prompt_version, so they never clash with the standard.

What not to log

Prompts, replies and tool arguments are where personal data lives: names, phone numbers, addresses, complaints. That is why the conventions make recording message content opt-in. For production, they suggest keeping content in a separate store with its own access rules and deletion schedule, and putting only a reference to it on the span.

  • Record what you need to debug, not everything you could.
  • Pseudonymise user IDs instead of storing phone numbers.
  • Mask before export, and set a retention period you actually enforce.
  • Know where traces go. A tracing service hosted abroad means personal data leaves the country.

That last point has legal weight. Chat logs from a WhatsApp bot are personal data. Nigeria's Data Protection Act 2023 requires collecting only what you need (section 24(1)(c)), keeping it no longer than necessary (section 24(1)(d)), and recording the legal basis when personal data leaves Nigeria (section 41). Kenya's Data Protection Act 2019 has similar rules on minimisation, retention and transfers.

In AfricaBefore you point your traces at a hosted service, find out which country it stores data in. If it's abroad, record the legal basis for the transfer, or self-host a tracing tool in-country.

The few numbers worth watching

Google's site reliability engineers start from four golden signals: latency, traffic, errors and saturation. For AI features, add cost and quality.

Always watch percentiles, not averages. In the opening example the p95, the time 95% of requests beat, is 12 seconds, while the average says 1.92. Alerts I'd set on any AI feature:

  • p95 latency, per operation, so a slow tool stands out from a slow model;
  • error rate, split by error.type;
  • replies cut off at the token limit, from gen_ai.response.finish_reasons;
  • cost per feature per day;
  • the failure rate from your online evals.

Sample production, then close the loop

You can't read every conversation, so sample. Husain and Shankar suggest keeping some random traces in every sample, plus the ones most likely to hide problems: errors, thumbs-down, unusually long conversations, handovers to a person. Run your code checks and judges on that sample. This is an online eval.

Track the failure rate with an interval, not a bare number. If 32 of 400 sampled replies fail a check, that's 8.0%, with a 95% Wilson interval of about 5.7% to 11.1%. A day that moves within that band isn't news.

Feedback buttons help, but Anthropic's engineers warn that user feedback is sparse and skewed towards the worst problems. Count retries, handovers and complaints too.

Then close the loop. Every new failure you find goes through error analysis: read the trace, name the failure. It becomes a test case in your offline set. That case runs on every change from then on. The test set grows from real problems rather than imagined ones, and the trace that exposed the problem is the evidence.

Tools you can run yourself

You don't need a vendor to start: spans written to a file and a short script that reports percentiles will take you a long way. When you want a UI, several tracing tools read OpenTelemetry and can be self-hosted. Their licences differ:

ToolLicenceSelf-host
LangfuseMIT, except its enterprise (ee) foldersYes
Arize PhoenixElastic License 2.0: source-available, not open sourceYes
Comet OpikApache 2.0Yes
OpenLLMetryApache 2.0A library built on OpenTelemetry; sends to any backend

I checked these licences in each repository in October 2026. Licences change, so check again before you build a product on one.

Learn it properly

This post is drawn from Track 8 of the AI Study Group, Evals and observability, which is free and ends with a one-file tracer and report you can run on a laptop.

Sources