Can you trust an LLM judge? Check it against your own labels first
How to test a judge model like the classifier it is, with true positive and true negative rates and Cohen's kappa against your own labels, which biases to design around, and how to tell a real improvement from luck.
Your judge model says 80% of last week's replies passed. Is that good news? You can't tell until you know how often the judge itself is wrong. A judge is just another model making predictions, and it can be wrong in ways that flatter your system.
LLM-as-judge is a useful tool. It can check thousands of replies a day for failures no regular expression could catch. But it only earns that job after you have measured it against people, the same way you'd test any classifier before trusting its output.
Use a judge only where code can't see the failure
A card number in a reply is a pattern: a regex finds it every time, for free. A refund promise the policy doesn't allow can be worded a hundred ways. That second kind of failure needs understanding, and that is where a judge is worth building.
Hamel Husain and Shreya Shankar's guidance on judges comes down to a few rules:
- One judge, one failure mode. A single judge asked to catch everything is hard to write and impossible to check.
- Pass or fail, not 1 to 10. A binary verdict forces you to define the line. Fine-grained scores bunch together and drift.
- A short critique before the verdict. One or two sentences of reasoning, then PASS or FAIL.
- Examples from your own labels, taken only from the slice you set aside for writing the prompt.
Judges have documented biases
Strong judges can agree with people about as often as people agree with each other. In Zheng and colleagues' study (NeurIPS 2023), GPT-4 agreed with expert raters on 85% of comparisons without a tie, while the experts agreed with each other on 81%. The same paper, and others, also found clear biases:
- Position. GPT-4 kept the same verdict when two answers swapped places only 65% of the time. Adding examples to the prompt raised that to 77.5%. In a separate study, just reordering the answers let Vicuna-13B beat ChatGPT on 66 of 80 questions, with ChatGPT as the judge.
- Length. Some judges preferred padded, repetitive answers more than 90% of the time.
- Maths without a reference. Grading maths answers on its own, GPT-4 got 14 of 20 wrong. Given the correct answer, it got 3 of 20 wrong.
- Language. Hada and colleagues (EACL 2024) found GPT-4-based evaluators leaning towards higher scores, especially in lower-resource languages and non-Latin scripts.
The practical response: judge pairs in both orders, give the judge a reference answer whenever one exists, and measure it separately for each language your users write in, with labels from a native speaker.
Label first, then split the labels
To measure a judge you need ground truth, and that means labelling replies yourself as pass or fail. Husain and Shankar suggest 100 to 200 labels per failure mode, with enough failures among them to matter. Then split them three ways:
- about 15% as examples inside the judge prompt;
- about 40% to improve the prompt against;
- about 45% kept back for one final test the judge has never seen.
If you tune the prompt on the same labels you report, the judge will look better than it is.
Measure TPR, TNR and kappa, not raw agreement
On the final test, compute three numbers:
- True positive rate (TPR): of the replies you passed, the share the judge passes.
- True negative rate (TNR): of the replies you failed, the share the judge catches. Failures are usually rare, so this is the one to watch.
- Cohen's kappa: agreement after removing the agreement you'd expect by chance.
def agreement(human, judge):
"""human and judge are lists of True (pass) / False (fail), in the same order."""
pairs = list(zip(human, judge))
n = len(pairs)
tp = sum(h and j for h, j in pairs) # you passed, judge passed
tn = sum(not h and not j for h, j in pairs) # you failed, judge failed
you_pass, judge_pass = sum(human), sum(judge)
observed = (tp + tn) / n
chance = (you_pass / n) * (judge_pass / n) + (1 - you_pass / n) * (1 - judge_pass / n)
return {
"agreement": observed,
"kappa": (observed - chance) / (1 - chance),
"tpr": tp / you_pass, # good replies the judge passes
"tnr": tn / (n - you_pass), # bad replies the judge catches
}
def corrected_pass_rate(judge_rate, tpr, tnr):
"""Rogan-Gladen: undo the judge's known mistakes. Needs tpr + tnr > 1."""
return (judge_rate + tnr - 1) / (tpr + tnr - 1)
# 100 labelled replies: 57 both pass, 3 you pass / judge fails,
# 12 you fail / judge passes, 28 both fail.
human = [True] * 57 + [True] * 3 + [False] * 12 + [False] * 28
judge = [True] * 57 + [False] * 3 + [True] * 12 + [False] * 28
m = agreement(human, judge)
print({k: round(v, 2) for k, v in m.items()})
# {'agreement': 0.85, 'kappa': 0.68, 'tpr': 0.95, 'tnr': 0.7}
print(round(corrected_pass_rate(0.80, m["tpr"], m["tnr"]), 2))
# 0.77
Why not just report "the judge agrees with me 85% of the time"? Because agreement hides the failures. Picture 100 replies of which only 5 are bad. A judge that passes almost everything can agree with you on 91 of them while catching just 1 of the 5 failures. Run that through the function above and you get a TNR of 20% and a kappa of about 0.13. The 91% looked fine. The other two numbers tell you the judge is close to useless for the job you hired it for.
Correct the pass rate it reports
A judge with known error rates reports a biased pass rate. The Rogan–Gladen correction, borrowed from medical screening, undoes it: true pass rate ≈ (judge's pass rate + TNR − 1) ÷ (TPR + TNR − 1). In the example, a judge that passed 80% of new replies, with TPR 95% and TNR 70%, implies a true pass rate of about 77%.
Treat that as an estimate, not a fact. TPR and TNR come from a limited number of labels, so put an interval on the result; the open-source judgy package does this. The correction also assumes the judge's prompt and model haven't changed and that new traffic looks like what you labelled. Re-check the judge whenever one of those changes.
And decide what the judge is for. One that lets three in ten bad replies through is fine for watching weekly trends, but too leaky to block a release on its own. Keep code checks and a regular human spot check alongside it.
Is the new prompt better, or just lucky?
Once you trust your graders, the next trap is reading too much into a small gain. Say version A passes 41 of 50 cases and version B passes 43: 82% against 86%. The 95% Wilson intervals are about 69% to 90% and 74% to 93%. They overlap almost completely.
Run both versions on the same cases and look only at the ones that changed. Suppose B fixed 4 and broke 2. If the change did nothing, each changed case is a coin toss between fixed and broken. McNemar's exact test asks how often luck gives a split at least that uneven. For 4 against 2, p ≈ 0.69: luck does that most of the time. Even 4 fixed and none broken gives p = 0.125, which never clears the usual 0.05 bar. By contrast, 9 fixed and 1 broken gives p ≈ 0.02, unlikely to be luck.
Your system adds noise of its own. In one experiment, 1,000 identical requests at temperature 0 to an open model returned 80 different completions, because servers batch requests differently under load (Thinking Machines Lab, 2025). So run the unchanged version twice first. Any gain smaller than the flips you see between those runs means nothing. Evan Miller's paper on error bars for evals covers paired comparisons and how many cases you need in more depth.
Whatever the p-value, read the broken cases. They tell you what the change cost.
Learn it properly
This post is drawn from Track 8 of the AI Study Group, Evals and observability, which is free and builds a judge, checks it against labels and compares versions with a script you keep.
Sources
- Using LLM-as-a-judge for evaluation: a complete guide (Husain): one failure mode per judge, critiques, checking against people
- AI evals: everything you need to know (Husain and Shankar): label counts and splits
- Judging LLM-as-a-judge with MT-Bench and Chatbot Arena (Zheng et al., NeurIPS 2023): agreement, position, length and maths results
- Large language models are not fair evaluators (Wang et al., 2023): order swapping
- Are large language model-based evaluators the solution to scaling up multilingual evaluation? (Hada et al., EACL 2024)
- judgy: corrected pass rates with intervals
- Adding error bars to evals (Miller, 2024): standard errors and paired comparisons
- Interval estimation for a binomial proportion (Brown, Cai and DasGupta, 2001): the Wilson interval
- Defeating nondeterminism in LLM inference (Thinking Machines Lab, 2025)