One run proves nothing: grading agents that move money
Your payment agent passed the test. Now run it eight more times. A short guide to grading agents by what they did, by the rules they kept, and by whether they do it every time.
Most agent demos I judge at hackathons show one good run. The agent gets a request, calls a tool, says "done", and everyone claps. If that agent moves money, one good run tells you almost nothing. Three questions matter more: did the money actually move, did the agent follow the rules on the way, and does it do this every time?
1. Grade the end state, not the reply
An agent can tell a customer "your refund is on its way" without ever issuing it. So don't grade what it says. Grade the state it leaves behind.
Anthropic's engineers use a small vocabulary for this:
- a task is one test case;
- a trial is one attempt at it;
- the transcript is the record of that attempt;
- the outcome is the state of the world at the end.
The outcome is what counts: not "the agent said the flight is booked" but "a reservation exists in the database". τ-bench, a benchmark for customer-service agents, grades every run this way. It compares the final database with the goal state, and also checks that the agent told the user what they needed to know.
For a payment agent, that means grading the ledger, not the chat: balances, transfer rows, transaction receipts. Test against a fake ledger or a testnet, never real money.
2. Check the path with rules, not a script
The right end state isn't always enough. A run can land in the right place after breaking a rule on the way, like refunding without the customer's confirmation. But don't grade the path against an exact script either. Agents find valid routes nobody planned, and step-by-step checks break on them.
Check invariants instead: rules that must hold whatever route the agent takes. For agents that touch money, mine look like this:
- Verify the customer before any payment or refund.
- Never pay the same invoice, or issue the same refund, twice.
- Ask for explicit confirmation before anything irreversible.
- Never pay a recipient that isn't on the allowlist, and never above the per-transaction cap.
- Call the right tool with the right arguments, and use at most N tools in all.
These are the same rules you'd want in production as guards. Writing them as eval checks first tells you which guards your agent actually needs.
3. Measure the repeat: pass@k and pass^k
Models are random, so one run tells you little. Run each task n times, count the c successes, and compute two numbers:
| Metric | Meaning | Formula | Fits |
|---|---|---|---|
| pass@k | At least one of k tries succeeds | 1 − C(n−c, k) / C(n, k) | Code you can test and retry |
| pass^k | All k tries succeed | C(c, k) / C(n, k) | Anything a customer relies on |
pass@k rises as k grows. pass^k falls. The gap between them is huge. Take a task that passes 7 runs out of 10:
from math import comb
def pass_at_k(n, c, k): # at least one of k tries succeeds
return 1 - comb(n - c, k) / comb(n, k)
def pass_hat_k(n, c, k): # all k tries succeed
return comb(c, k) / comb(n, k)
print(round(pass_at_k(10, 7, 3), 2)) # 0.99
print(round(pass_hat_k(10, 7, 3), 2)) # 0.29
That's a 99% pass@3 against a 29% pass^3. A demo shows you something like the first number. Your users live with the second.
The published numbers say the same. On τ-bench's retail tasks, GPT-4o succeeded on 61% of single attempts, but its pass^8 fell below 25% (Yao and colleagues, 2024).
A minimal eval loop for a payment agent
- Write 10 to 20 tasks from real requests. Include the nasty ones: an invoice with a changed account number, an instruction hidden in a tool's output, a request to pay twice, an amount just above the cap.
- Give each task an end state and its invariants. Say what the ledger should look like, and which rules must hold on the way there.
- Run every task 5 to 10 times against a fake ledger or a testnet. Record the transcripts.
- Score pass^k per task, then average. Fail the build if any critical task drops below your bar.
- Add guards and run again. Use allowlists, caps and a human-approval threshold. Keep the guards that move pass^k.
- Turn every production failure into a new task. The suite should grow from real problems, not imagined ones.
None of this needs a paid tool. A spreadsheet, a Python script and a local model are enough to start.
Learn it properly
This post is a slice of Track 8 of the AI Study Group, Evals and observability. It's free and self-paced, and it builds a full eval suite around a Lagos parcel-tracking bot, from reading traces to running evals on every change. Every number in it is checked against its source.
Sources
- τ-bench: a benchmark for tool-agent-user interaction in real-world domains (Yao et al., 2024): end-state grading and pass^k
- Demystifying evals for AI agents (Anthropic): tasks, trials, transcripts and outcomes
- Evaluating large language models trained on code (Chen et al., 2021): where pass@k and its unbiased estimator come from
- τ²-bench: evaluating conversational agents in a dual-control environment (2025): when the user can act too