Evaluation lesson 04 · Audit grounded claims
Shared evaluation case · deterministic fictional data
Northstar Support Copilot. A fictional ecommerce support copilot retrieves versioned policies and synthetic customer records, drafts replies, and proposes sandboxed account actions. It cannot send or mutate anything without current authorization and exact human approval.
01 · Smallest useful mechanism
A groundedness evaluation starts with gold evidence or clearly defined authority rules. It measures whether the right evidence was retrievable, whether the system selected it, and whether each material claim follows from it. Unsupported questions should produce calibrated abstention or escalation.
Citation presence is not evidence quality; authority, freshness, relevance, and entailment all matter.
02 · Experiment
Deterministic fictional fixture
The audit grounded claims workbench uses versioned, inspectable teaching data. It does not claim to reproduce live customer traffic or model behavior.
Evidence workbench
Choose the checks needed before calling a policy answer grounded.
Fixture northstar-grounding-v1
03 · Decision artifact
Use synthetic or properly governed data. Keep critical failures visible instead of compressing them into one score.
Separate the retrieval-quality claim from the answer-faithfulness claim.
Build a pack with gold passages, stale distractors, conflicts, and unsupported questions.
Export a claim ledger with exact spans, versions, dispositions, and required repairs.
Your learning artifact
Stored only in this browser. No account required; course reset does not delete it.
04 · Check your understanding
Theory notes and sources
A perfectly faithful answer to an obsolete refund policy is still wrong.
Common mistake: If every sentence has a citation, the response is grounded. A citation can be stale, irrelevant, from another customer, or fail to entail the claim.
Transfer exercise: Label each material claim supported, contradicted, stale, irrelevant, or unverifiable.
Next: Evidence-backed words are only half the product; the next lesson evaluates whether tool proposals and real state changes obey policy.