Buying Human Judgment
Automated scores ultimately rest on human judgment. Knowing when to spend on humans — and how to tell if that spend is producing reliable labels — is a budget and quality decision that lands on your desk.
Humans are slow and expensive, so spend them where they're irreplaceable: building the trusted "golden" set, validating the LLM judge, and the high-stakes or subjective calls automation can't be trusted on. The make-or-break detail is the rubric: vague instructions produce noisy, useless labels, no matter how much you pay. The health check is inter-annotator agreement — do two humans given the same case agree? If not, the rubric is ambiguous (not the people), and every downstream number built on those labels is shaky.
The labeling budget request
Your team asks for budget to have humans label 50,000 outputs, and mentions the two reviewers "often disagree." What do you fund, and what do you fix first?
Don't fund 50,000 labels yet — the disagreement is the headline. If two reviewers often disagree, the rubric is ambiguous, and you'd be buying 50,000 noisy labels at full price. Fix the rubric first (sharper criteria, examples for each judgment), confirm agreement improves, then label a focused, high-value set (a golden set + judge-validation set) rather than everything. Pair humans with a validated LLM judge to scale: humans set and audit the standard; the judge applies it cheaply.
Ravel: paying for labels that weren't ground truth
Ravel paid annotators to label examples that would serve as the "truth" their evals were measured against — then couldn't understand why nothing lined up. The problem was label quality: three annotators working from a vague one-line guideline agreed with each other only 64% of the time. They'd bought three people's differing opinions, not a truth, so every downstream number rested on sand.
The fix was process, not spend: a written rubric with examples, a calibration round on shared cases, and a senior reviewer to adjudicate disagreements. Agreement rose to 89%, and the labels became a yardstick worth trusting. The leadership lesson when you buy human judgment: the cost isn't just labeler hours — it's the guidelines, calibration, and adjudication that make the labels reliable. Cheap, unmanaged labeling produces expensive, misleading evals.
This chapter: Meridian has two experienced support leads label 120 of Remi's refund replies against a clear rubric, adjudicating disagreements to reach 90% agreement. That carefully-built human-labeled set is what exposed the Chapter-5 judge as only 56% right on refunds. The COO's takeaway: a little well-managed human labeling on the high-stakes slice is worth more than a lot of sloppy labeling everywhere.
Quiz · Chapter 6
- Humans are best spent on:
- If two annotators frequently disagree, the usual cause is:
- Before funding 50,000 hand labels, you should:
- Inter-annotator agreement tells you: