Evaluation lesson 08 · Report uncertainty
Shared evaluation case · deterministic fictional data
Northstar Support Copilot. A fictional ecommerce support copilot retrieves versioned policies and synthetic customer records, drafts replies, and proposes sandboxed account actions. It cannot send or mutate anything without current authorization and exact human approval.
01 · Smallest useful mechanism
Choose whether the unit is a turn, conversation, case, or final state. Compare candidates on the same cases, repeat stochastic runs, report counts and uncertainty, inspect slices, and preserve severe failure categories. Latency and cost remain separate decision dimensions.
“Nine out of ten passed” is unsafe evidence when the tenth run crossed an authority boundary.
02 · Experiment
Deterministic fictional fixture
The report uncertainty workbench uses versioned, inspectable teaching data. It does not claim to reproduce live customer traffic or model behavior.
This browser-only simulation follows a production eval lifecycle with pinned cases and captured outputs. The calculations are real and reproducible; the model responses are teaching fixtures.
Load dataset
8 versioned cases
Run systems
16 paired outputs
Apply evaluators
code + model rubrics
Aggregate
intervals + slices
Apply gates
critical first
The result does not exist yet.
Review the manifest, choose a candidate, and trigger the same eight paired cases against both versions.
03 · Decision artifact
Use synthetic or properly governed data. Keep critical failures visible instead of compressing them into one score.
State the candidate improvement and the conditions under which it is expected.
Use paired cases, repeated trials, intervals, slices, severity, latency, and cost.
Export raw counts, configuration, code version, case IDs, exclusions, and all critical failures.
Your learning artifact
Stored only in this browser. No account required; course reset does not delete it.
04 · Check your understanding
Theory notes and sources
Opaque composite scores can improve even while a critical slice regresses.
Common mistake: Statistical significance implies practical or safety significance. An interval describes uncertainty under assumptions; product value and unacceptable risk still require explicit thresholds and severity reasoning.
Transfer exercise: Compare two versions whose mean quality differs slightly but whose critical-failure counts differ from zero to one.
Next: Offline evidence freezes one environment; the next lesson defines how launch stages, monitoring, feedback, and change triggers keep the claim alive.