The Course

Domain 4 · 16% of the exam

Hope is not
a test framework.

Every factory has an inspection line — except, somehow, most AI systems. This episode is about knowing, not believing: metrics, eval sets, A/B tests, and the diagnostic instincts that find the real fault in minutes.

Objectives

What you’ll be able to do

  1. Define metrics: accuracy, latency, cost, safety, security
  2. Design eval datasets and mixed-method test frameworks
  3. Run A/B tests and iterate with evidence
  4. Diagnose prompt failures, hallucinations, model mismatch
  5. Optimize tokens, latency, and cost–performance trade-offs
  6. Monitor systems with logging and observability

The story

The Friday deploy

Theo improved a prompt on a Friday. He tried it on three claims — "all three looked better!" — and shipped it. By Tuesday, the water-damage team was in revolt: approvals that used to sail through were being flagged, and nobody could say why. "How do you know the new prompt is better?" Meera asked. "It felt better." She wrote two words on the whiteboard: felt and measured, and drew a wall between them.

They built the inspection line that weekend. Two hundred claims with known-correct outcomes — the golden set, drawn from real adjuster decisions, including the ugly edge cases. Automated checks for format and policy citations. An LLM judge scoring reasoning quality, spot-audited by humans. Every prompt change now walked the line before it touched production, and A/B tests settled arguments that used to be shouting matches.

The next "better" prompt lost to the old one, 61 to 74, on the golden set. It never shipped. Nobody argued. That silence, Meera told Theo, is what evaluation buys: the end of vibes as a deployment strategy.

Metaphor map

Images that stick

ConceptImageWhy it sticks
EvaluationThe inspection lineNothing reaches the loading dock without passing the stations. Not the flagship product, not the "small fix".
Golden setCrash-test dummiesReal crashes, recreated on demand. You don’t wait for a customer to find the failure at highway speed.
A/B testThe blind taste testTwo recipes, unlabeled, judged by volume. The chef’s confidence is not an ingredient.
HallucinationThe confident tour guideFluent, charming, and describing a building that isn’t there. Confidence is not a signal of truth.
Latency budgetTraffic signalsEvery stage owns a slice of the journey. Find the red light before blaming the road.
Cost meterThe fuel gaugeWatched per trip, not per quarter. A leak found at the pump is cheap; found on the highway, it’s a crisis.

Diagnosis console

Diagnosis console

The exam loves symptom → most-likely-cause questions. Pick a symptom; learn where a working architect looks first.

Retrieval / indexing failure

The only thing that changed is the corpus. A broken re-index, mismatched embeddings, or stale chunks are feeding the model poor context — and a model given wrong context answers wrongly with full confidence.

Symptom
Confident but wrong answers began right after a document refresh; latency and model unchanged
Do first
the retrieved chunks in a failing trace — are they relevant, current, correctly re-indexed?
Not this
model weights, temperature, or the prompt — none of them changed.

Context overflow / eviction

As history grows, early content is truncated, summarized away, or simply drowned. The instruction isn’t being disobeyed — it’s no longer effectively present.

Symptom
The assistant ignores instructions that were given early in long conversations
Do first
the assembled request near failure: is the instruction still there, and how far from the end?
Not this
model quality — the same model obeys the same instruction in short sessions.

Prompt/format brittleness

Intermittent structure failures point to weak format constraints — freeform format instructions instead of enforced structured output, or examples that conflict with the schema.

Symptom
Answers are factually fine but the JSON format breaks intermittently under load
Do first
failing outputs side-by-side with the format spec; tighten with structured-output enforcement and consistent examples.
Not this
retrieval or infrastructure — the content is correct; only the shape wobbles.

Perceived latency (no streaming)

Total generation time may be acceptable, but nothing appears until the full response is done. Streaming and a cached prefix move the first visible token from seconds to near-instant.

Symptom
Quality is fine, but users complain the assistant "hangs" before answering
Do first
time-to-first-token vs total time in traces; enable streaming, cache the static prefix.
Not this
model capability — this is a delivery problem, not an intelligence problem.

Model–task mismatch

The downgraded tier handles surface tasks but lacks depth for multi-step reasoning. Aggregate metrics can hide this — the failures cluster in the hard slice of traffic.

Symptom
A cheaper model was swapped in; simple queries are fine but multi-step cases quietly degrade
Do first
eval results segmented by difficulty; route complex cases to the capable tier or revert.
Not this
prompts or retrieval — they didn’t change; the reasoning budget did.

Optimization lab

Optimization lab

A RAG assistant: 3.5s to first token, 9s total, quality at baseline. Toggle optimizations and watch what each one really buys — and costs. Illustrative numbers; the trade-off shapes are the lesson.

Streaming

Show tokens as they generate — first words appear at TTFT, not at the end

Prompt caching

Stable prefix precomputed — cuts TTFT and input cost, zero quality impact

Trim retrieval (top-k ↓)

Fewer, better chunks — cheaper and faster, small quality risk

Smaller model tier

Much cheaper and faster — the only lever that spends quality directly

In production

A product-Q&A bot, measured into shape

In production · retail

A product-Q&A bot, measured into shape

BASELINE
1,000-question golden set from real chats; code checks for price/stock facts, LLM judge for helpfulness, weekly human audit of 50.
DIAGNOSE
Post-catalog-refresh accuracy dips traced to stale index chunks in minutes — the trace showed retrieval, not the model, had changed.
A/B
New answer format tested on 10% of traffic for two weeks: +6 points helpfulness, no accuracy change, then rolled to 100%.
OPTIMIZE
Caching + streaming cut perceived wait dramatically at zero quality cost; a model downgrade was tested, failed the golden set, and was rejected — with evidence.

Storyboard

The film, shot by shot

01
8s

Every factory has one. The line nothing skips.

Cameralow tracking shot down a gleaming inspection corridor, stations receding to vanishing point
Motionproducts glide station to station; stamps fall: PASS, PASS, PASS
Soundconveyor rhythm, pneumatic stamp, room tone
02
10s

Except here. Here, someone tried three examples on a Friday… and felt good.

Camerawhip-pan to a dim side door labeled PRODUCTION, propped open with a coffee cup
Motiona glowing prompt-scroll sneaks past the line; the corridor lights flicker uneasily
Soundrecord scratch of the conveyor, door creak, distant alarm
03
12s

By Tuesday, the complaints. Fluent answers. Confident answers. Wrong answers.

Camerasplit screen: cheerful chat bubbles left, a red defect graph climbing right
Motionbubbles sprout tiny cracks as they land; the graph line snakes upward like smoke
Soundmessage pings curdling into minor key, graph tick like a Geiger counter
04
12s

So we build the dummies. Two hundred real cases, the ugly ones included, answers known in advance.

Cameraoverhead god-shot of a test hall filling with crash-test figures on numbered plinths
Motioneach dummy assembles from archived tickets; plinth numbers count to 200
Soundmechanical assembly clicks, numbers whirring like a departure board
05
10s

Now every change walks the line. Facts checked by code. Judgment scored by a judge. The judge audited by humans.

Camerathree-station dolly: scanner gate, robed judge figure, human with a clipboard
Motiongreen lasers grade format; the judge weighs answer-scales; the human corrects the judge’s stamp once
Soundscanner sweep, scale creak, pencil scratch
06
10s

The new prompt scored sixty-one against seventy-four. It never shipped. Nobody argued. That silence is the product.

Cameraslow push on a scoreboard in an empty hall, lights shutting down row by row
Motion61 vs 74 burns on the board; the losing scroll files itself into a drawer marked NEXT ITERATION
Soundsingle chime, drawer slide, satisfying silence; title card

Flashcards

Recall drill

Select a card to reveal the answer.

Quiz

Exam drill

Question 1

A team wants to replace its prompt with a "clearly better" rewrite. They tried it on four hand-picked examples and all four looked good. What must happen before this ships?

Question 2

You must evaluate an assistant whose answers require judgment (tone, reasoning quality) as well as factual precision (prices, dates). Which evaluation design fits best?

Question 3

An assistant meets its quality bar but misses its cost target by 40%. Traces show a 7k-token static prefix on every call and full responses generated before anything is shown. What should you try FIRST?

Revision

Close your eyes and remember

Felt is not measured. The golden set is scar tissue. Five metrics or it didn't happen. Symptom, then trace, then cause. And the free lunches come first: stream, cache, then — only with proof — touch the model.

Domain 4 · 16% of your exam

Next: Domain 5