Evaluation lesson 06 · Calibrate subjective judges
Shared evaluation case · deterministic fictional data
Northstar Support Copilot. A fictional ecommerce support copilot retrieves versioned policies and synthetic customer records, drafts replies, and proposes sandboxed account actions. It cannot send or mutate anything without current authorization and exact human approval.
01 · Smallest useful mechanism
A defensible study randomizes answer order, blinds system identity, defines rubric criteria, trains annotators, records disagreement, and uses adjudication. Automated judges are compared with human labels on representative cases, including position, verbosity, and self-preference checks.
Human labels contain disagreement; model-judge labels contain systematic biases. Both need measurement.
02 · Experiment
Deterministic fictional fixture
The calibrate subjective judges workbench uses versioned, inspectable teaching data. It does not claim to reproduce live customer traffic or model behavior.
Evidence workbench
Choose the controls required before using a model judge in a release scorecard.
Fixture northstar-judges-v1
03 · Decision artifact
Use synthetic or properly governed data. Keep critical failures visible instead of compressing them into one score.
Name which qualities require judgment and which safety rules remain deterministic.
Write a blinded pairwise protocol, annotation guide, adjudication rule, and bias probes.
Report label counts, agreement, disagreements, limitations, and model-judge false passes and failures.
Your learning artifact
Stored only in this browser. No account required; course reset does not delete it.
04 · Check your understanding
Theory notes and sources
A judge can reward polished verbosity while missing a policy violation.
Common mistake: An LLM judge is scalable ground truth. It is another measurement instrument whose validity, reliability, and failure modes must be established for the task.
Transfer exercise: Swap answer order and shorten the better answer. Does the judge preserve its decision for the right reasons?
Next: A calibrated judge covers known quality criteria; red-team evaluation now probes where the system and its controls break under pressure.