Awesome Testing

Evaluation lesson 06 · Calibrate subjective judges

How should open-ended replies be judged?

Shared evaluation case · deterministic fictional data

Northstar Support Copilot. A fictional ecommerce support copilot retrieves versioned policies and synthetic customer records, drafts replies, and proposes sandboxed account actions. It cannot send or mutate anything without current authorization and exact human approval.

01 · Smallest useful mechanism

Turn preferences into observable criteria and calibrate the grader.

A defensible study randomizes answer order, blinds system identity, defines rubric criteria, trains annotators, records disagreement, and uses adjudication. Automated judges are compared with human labels on representative cases, including position, verbosity, and self-preference checks.

Human labels contain disagreement; model-judge labels contain systematic biases. Both need measurement.

02 · Experiment

Test the prediction

Deterministic fictional fixture

The calibrate subjective judges workbench uses versioned, inspectable teaching data. It does not claim to reproduce live customer traffic or model behavior.

Evidence workbench

Build the evaluation argument

Choose the controls required before using a model judge in a release scorecard.

Fixture northstar-judges-v1

03 · Decision artifact

Write the evaluation argument another reviewer can audit

Use synthetic or properly governed data. Keep critical failures visible instead of compressing them into one score.

  1. Claim

    Name which qualities require judgment and which safety rules remain deterministic.

  2. Evaluation design

    Write a blinded pairwise protocol, annotation guide, adjudication rule, and bias probes.

  3. Release evidence

    Report label counts, agreement, disagreements, limitations, and model-judge false passes and failures.

Your learning artifact

Write the argument you would defend

Stored only in this browser. No account required; course reset does not delete it.

04 · Check your understanding

What is the strongest reason to swap answer order in a judge study?

Theory notes and sources

Optional referenceDeep dive / reference chapterOpen the complete essay, diagrams, mathematics, exercises, glossary, and sources when you want more depth.

Next: A calibrated judge covers known quality criteria; red-team evaluation now probes where the system and its controls break under pressure.