Start from real conversations: evals built from what users actually sent
A practical method for reading traces, naming what goes wrong, counting it, and turning the worst failures into a small test set that looks like your real traffic.
You shipped an AI feature after trying a handful of messages. The replies looked fine. A month later complaints are rising, every prompt fix seems to break something else, and someone suggests paying for a bigger model. Nobody can say what is actually going wrong.
The tempting move is to reach for a benchmark. Public leaderboards are evals too, but they test models in general, mostly in English. They help you choose a model to start with. They can't tell you whether your support bot quotes the right refund policy to a customer writing in Pidgin.
More spot checks won't save you either. If a failure hits 1 user in 10, five random conversations show no sign of it 59% of the time (0.9 to the power of 5). At 1 in 20, it stays hidden 77% of the time.
What works is slower and less glamorous: read what your system actually did, write down what went wrong, count it, and test for the failures that matter. Hamel Husain and Shreya Shankar, who teach this to engineering teams, call it error analysis. Here is how I do it.
Read traces, and note the first thing that broke
A trace is the full record of one interaction: the user's messages, your system's replies, and everything in between, such as tool calls and their results, retrieved documents and each step an agent took. You don't need a platform to start. A spreadsheet with one row per conversation is enough.
Read what the user saw first, then the details. For each trace, write down the first thing that went wrong. Later problems usually follow from it. If a parcel lookup returned nothing and the bot then invented a delivery date, and later apologised with a second invented date, the note is "invented a status when the lookup failed". The apology is a symptom.
Open coding: plain notes, no categories yet
Go through the traces one by one. If a trace is fine, mark it as a pass. If not, write one short note in plain words: "said out for delivery, but the lookup returned nothing", "answered in English to a Pidgin message". Researchers call this open coding.
- Write what you see, not why it happened. Causes come later.
- Note problems that aren't the model's fault, such as a missing link to the price list.
- Do the first 30 or more yourself. Husain and Shankar advise against handing this first pass to an LLM. This is where you learn what "good" means for your product.
- Don't write the rules first. In a study of people building evaluators, participants needed criteria to grade outputs, yet only found their criteria by grading outputs. The authors named this criteria drift (Shankar and colleagues, 2024).
Pick one person whose judgement defines "good", usually the domain expert: for a support bot, the head of support. Husain and Shankar call this role the benevolent dictator. If several people label, have each of them label the same 20 traces separately first, and check how often they agree.
Axial coding: name the failure modes, then count
Once you have notes on a few dozen failures, group similar notes into failure modes. This step is called axial coding. An LLM can suggest groupings here, but you decide.
The test of a good name is whether a yes-or-no check could apply it. "Reliability issues" fails that test. "Invents a parcel status when the lookup fails" passes it. Here is what that looks like for a parcel-delivery support bot:
| Vague bucket | Testable failure mode | How you might check it |
|---|---|---|
| Reliability issues | Invents a parcel status when the lookup fails | Lookup returned nothing, reply states a status |
| Language problems | Replies in English to a Pidgin message | Compare the language of message and reply |
| Policy issues | Promises a refund the policy doesn't allow | A judge model checked against your labels |
| Escalation | No handover when the customer asks for a person | Did the handover tool run? |
| Formatting | Too long to read on a phone | A word or character limit in code |
Then count. A pivot table works. So do twenty lines of Python with no dependencies:
import csv
import io
from collections import Counter
# One row per trace you read. Leave "category" empty for a pass.
REVIEWS = """trace_id,verdict,category
t001,pass,
t002,fail,invents a parcel status when the lookup fails
t003,fail,replies in English to a Pidgin message
t004,pass,
t005,fail,promises a refund the policy doesn't allow
t006,fail,invents a parcel status when the lookup fails
t007,pass,
t008,fail,no handover when asked for a person
"""
BLOCKING = {"promises a refund the policy doesn't allow"}
def tally(rows):
rows = list(rows)
fails = [r["category"].strip() for r in rows if r["verdict"] == "fail"]
print(f"{len(rows)} traces read, {len(fails)} failed ({len(fails) / len(rows):.0%}).\n")
running = 0
for category, count in Counter(fails).most_common():
running += count
flag = " [BLOCKS RELEASE]" if category in BLOCKING else ""
print(f"{count:4} {count / len(fails):4.0%} of failures"
f" (running {running / len(fails):4.0%}) {category}{flag}")
# Use open("reviews.csv", newline="", encoding="utf-8") for your own file.
tally(csv.DictReader(io.StringIO(REVIEWS)))
The running total shows how few categories cause most of the pain. Fix the most common serious ones first. Keep failures that must block a release, such as a promise that costs money or leaked personal data, flagged on their own, however rare they are.
This pays off. In a case Husain describes, an apartment-leasing assistant's team found that three issues caused over 60% of its problems. One of them: requests with relative dates, such as "two weeks from now", failed about two-thirds of the time. After targeted fixes and tests, the team reported success on those requests rising from 33% to 95%.
When to stop reading
Keep going until new traces stop teaching you anything: no new categories, no changes to your definitions. Researchers call this saturation. Husain and Shankar suggest at least 100 traces, and the arithmetic agrees. The chance of seeing a failure at least once in n random traces is 1 − (1 − rate)n:
- A 1-in-50 failure appears in 30 traces only 45% of the time, and in 100 traces 87% of the time.
- To be 95% sure of seeing a 1-in-100 failure even once, you need about 299 random traces.
So for rare failures, stop sampling at random and search on purpose: conversations that ended in a complaint, unusually long ones, and ones where a tool failed.
Turn the failures into a small test set
Your notes now point at the cases worth testing. Collect them from the most trustworthy sources first:
- Real conversations, especially the failures you just found.
- Support tickets and complaints: each one is a case your system already got wrong.
- The checks you already do by hand before a release.
- Drafted cases, to fill gaps the real data doesn't cover yet.
- Edge cases: a one-word message, typos, two questions in one, an attempt to trick the bot, a question it can't answer.
For drafts, list the dimensions along which real requests differ: who is writing, what they need, which language, whether they gave a tracking number. One value from each is a tuple. Three customer types, five needs, two languages and two tracking states make 60 combinations. You don't need all of them, just every value covered plus the combinations where failures cluster. Husain and Shankar suggest writing about 20 tuples by hand, then letting a model turn them into realistic messages. Generate inputs only, run them through your real system, and read the replies.
What each case needs
Each case needs an input, a plain statement of what a good reply must and must never do, a reference answer where one exists, and tags for where it came from and what it covers. Grade the outcome, not the exact wording: many different replies can be right. Balance the set, too: include customers who should be handed to a person and customers who shouldn't.
You don't need hundreds. Anthropic's engineers say 20 to 50 tasks drawn from real failures make a good start. Hold a slice back that you never look at while tuning prompts, or your score will climb while users notice nothing. Give the set a version number and a one-line changelog, because scores on different versions can't be compared.
One more thing: read and write in your users' language. Husain and Shankar warn that machine translation loses meaning and cultural cues, and that model-written drafts are least reliable in lower-resource languages. When researchers built INJONGO, a test of requests in 16 African languages, native speakers wrote the requests themselves instead of translating English ones. If you don't speak Pidgin, Hausa or Twi, find a reviewer who does.
Learn it properly
This post is drawn from Track 8 of the AI Study Group, Evals and observability, which is free and walks through error analysis, test sets and graders around one worked example.
Sources
- Why is error analysis so important, and how is it performed? (Husain and Shankar): open and axial coding, reading 100 traces
- A field guide to rapidly improving AI products (Husain): the apartment-leasing case
- Who validates the validators? (Shankar et al., 2024): criteria drift
- What is the best approach for generating synthetic data? (Husain and Shankar): dimensions and tuples
- Demystifying evals for AI agents (Anthropic): starting with 20 to 50 real tasks
- INJONGO: a multicultural intent detection and slot-filling dataset (Yu et al., ACL 2025): requests written by native speakers