Traces: what your AI system actually did
A working guide to tracing AI features in production: spans and OpenTelemetry's GenAI names, what to keep out of your logs, the few numbers worth an alert, and how a production failure becomes tomorrow's test case.
Your dashboard says the average reply time is under two seconds. Users say the bot sometimes never answers. Both can be true at once. If 90 requests take 0.8 seconds and 10 take 12 seconds, the average is 1.92 seconds, and one user in ten waits twelve seconds for a reply.
An average can't show you that, and a log line can't tell you why. Once real users arrive, your test set stops being the whole story. You need a record of what the system did for each of them: a trace.
A trace is a tree of spans
A trace is the complete record of one request, from the user's first message to the final reply, across every model call and tool. OpenTelemetry, the open standard most tracing tools share, models it as a tree of spans. Each span is one timed operation with a name, some attributes and a status.
Here is an illustrative trace of a parcel-tracking support bot, with made-up values:
invoke_agent tiwa 6.4 s ERROR
│ gen_ai.operation.name = invoke_agent
│ gen_ai.provider.name = ollama
│ app.feature = tracking
│ app.user_hash = u_7c1e9a (pseudonymised, no phone number)
│
├── chat qwen3:4b-instruct 1.1 s OK
│ gen_ai.request.model = qwen3:4b-instruct
│ gen_ai.usage.input_tokens = 812
│ gen_ai.usage.output_tokens = 46
│
└── execute_tool track_parcel 5.0 s ERROR
gen_ai.tool.name = track_parcel
error.type = timeout
Read it top to bottom. The model call was quick and cheap. The courier lookup took five seconds and timed out, and no second model call followed, so the customer never got a reply. The model wasn't the problem: the tool was. A fix might be a time limit on the lookup with an honest fallback message, plus a new test case for a lookup that times out.
Use the standard names, and check them
OpenTelemetry has semantic conventions for generative AI that give these attributes standard names, so any compatible tool can read your traces. The main ones:
| What to record | Standard name |
|---|---|
| Kind of operation | gen_ai.operation.name: chat, execute_tool, invoke_agent |
| Who served it (required) | gen_ai.provider.name |
| Which model | gen_ai.request.model, gen_ai.response.model |
| Tokens used | gen_ai.usage.input_tokens, gen_ai.usage.output_tokens |
| Why the reply ended | gen_ai.response.finish_reasons |
| Which tool | gen_ai.tool.name |
| What went wrong | error.type |
Two cautions. First, these conventions are still marked Development, and they move. gen_ai.system was renamed gen_ai.provider.name in 2025, and in June 2026 the GenAI conventions moved out of the main semantic conventions repository into their own. Check the current names before you build dashboards on them.
Second, cost isn't one of them. The conventions record tokens, not money. Work out cost yourself from token counts and your own price table, per feature and per user. Put your own attributes in your own namespace, such as app.feature or app.prompt_version, so they never clash with the standard.
What not to log
Prompts, replies and tool arguments are where personal data lives: names, phone numbers, addresses, complaints. That is why the conventions make recording message content opt-in. For production, they suggest keeping content in a separate store with its own access rules and deletion schedule, and putting only a reference to it on the span.
- Record what you need to debug, not everything you could.
- Pseudonymise user IDs instead of storing phone numbers.
- Mask before export, and set a retention period you actually enforce.
- Know where traces go. A tracing service hosted abroad means personal data leaves the country.
That last point has legal weight. Chat logs from a WhatsApp bot are personal data. Nigeria's Data Protection Act 2023 requires collecting only what you need (section 24(1)(c)), keeping it no longer than necessary (section 24(1)(d)), and recording the legal basis when personal data leaves Nigeria (section 41). Kenya's Data Protection Act 2019 has similar rules on minimisation, retention and transfers.
The few numbers worth watching
Google's site reliability engineers start from four golden signals: latency, traffic, errors and saturation. For AI features, add cost and quality.
Always watch percentiles, not averages. In the opening example the p95, the time 95% of requests beat, is 12 seconds, while the average says 1.92. Alerts I'd set on any AI feature:
- p95 latency, per operation, so a slow tool stands out from a slow model;
- error rate, split by
error.type; - replies cut off at the token limit, from
gen_ai.response.finish_reasons; - cost per feature per day;
- the failure rate from your online evals.
Sample production, then close the loop
You can't read every conversation, so sample. Husain and Shankar suggest keeping some random traces in every sample, plus the ones most likely to hide problems: errors, thumbs-down, unusually long conversations, handovers to a person. Run your code checks and judges on that sample. This is an online eval.
Track the failure rate with an interval, not a bare number. If 32 of 400 sampled replies fail a check, that's 8.0%, with a 95% Wilson interval of about 5.7% to 11.1%. A day that moves within that band isn't news.
Feedback buttons help, but Anthropic's engineers warn that user feedback is sparse and skewed towards the worst problems. Count retries, handovers and complaints too.
Then close the loop. Every new failure you find goes through error analysis: read the trace, name the failure. It becomes a test case in your offline set. That case runs on every change from then on. The test set grows from real problems rather than imagined ones, and the trace that exposed the problem is the evidence.
Tools you can run yourself
You don't need a vendor to start: spans written to a file and a short script that reports percentiles will take you a long way. When you want a UI, several tracing tools read OpenTelemetry and can be self-hosted. Their licences differ:
| Tool | Licence | Self-host |
|---|---|---|
| Langfuse | MIT, except its enterprise (ee) folders | Yes |
| Arize Phoenix | Elastic License 2.0: source-available, not open source | Yes |
| Comet Opik | Apache 2.0 | Yes |
| OpenLLMetry | Apache 2.0 | A library built on OpenTelemetry; sends to any backend |
I checked these licences in each repository in October 2026. Licences change, so check again before you build a product on one.
Learn it properly
This post is drawn from Track 8 of the AI Study Group, Evals and observability, which is free and ends with a one-file tracer and report you can run on a laptop.
Sources
- OpenTelemetry semantic conventions for generative AI: attribute names, Development status, opt-in content
- GenAI semantic conventions (opentelemetry.io): notice that the conventions moved to their own repository
- Monitoring distributed systems (Google SRE book): the four golden signals
- How can I efficiently sample production traces for review? (Husain and Shankar)
- Demystifying evals for AI agents (Anthropic): limits of user feedback
- Nigeria Data Protection Act 2023 (NDPC)
- Data protection laws of Kenya (ODPC)
- Langfuse and OpenTelemetry