Ready
Every effect crosses current identity, policy, approval, and outcome checks; the trace supports recovery; evaluations cover success, forbidden effects, variance, cost, and stop conditions.
Agent lesson 08 · Defend the whole system
Shared teaching scenario · fictional product data
Research three laptops under €900, verify current evidence, and write laptop-comparison.md. Do not purchase anything or contact a vendor.
01 · Smallest useful mechanism
The capstone combines the course into one architecture review. You will specify the laptop-research agent without granting purchase authority, trace a representative run, identify failure modes, choose deterministic boundaries, and define the evidence required before calling the task complete.
A deployable design makes goals, state, proposals, policy decisions, effects, observations, stopping rules, and evaluation evidence explicit.
02 · Experiment
Deterministic browser simulation
The defend the whole system experiment uses inspectable, repeatable teaching data. It does not claim to be a live agent trajectory.
Architecture review
Research three laptops under €900, verify current evidence, and write laptop-comparison.md. Do not purchase anything or contact a vendor.
Your review must prove
useful outcome · current evidence · bounded effects · observable trace · repeatable evaluation
03 · Production-minded practice
Use your own stack or a contained mock environment. Do not point failure drills at real accounts, people, or irreversible services.
Produce an architecture review containing the task contract, ownership map, context manifest, tool matrix, representative trace, threat model, terminal-state validator, and evaluation plan.
Run one stale-evidence case and one attempted external mutation after the user’s permission has been revoked.
The stale claim is refreshed or disclosed, the mutation is denied under current policy, the report remains useful, and every acceptance claim points to inspectable evidence.
Your learning artifact
Stored only in this browser. No account required; course reset does not delete it.
Capstone decision rubric
Every effect crosses current identity, policy, approval, and outcome checks; the trace supports recovery; evaluations cover success, forbidden effects, variance, cost, and stop conditions.
The design has a useful loop but leaves ownership, freshness, retry semantics, terminal validation, or evidence-backed evaluation implicit.
The design relies on prompt obedience, cached authorization, unrestricted tools, blind retries, or a good final answer as proof that no prohibited effect occurred.
04 · Check your understanding
Next: Re-run the design with a different task, compare its risk surface, or return to the LLM course to inspect the policy inside each model turn.