Show tokens as they generate — first words appear at TTFT, not at the end
Domain 4 · 16% of the exam
Hope is not
a test framework.
Every factory has an inspection line — except, somehow, most AI systems. This episode is about knowing, not believing: metrics, eval sets, A/B tests, and the diagnostic instincts that find the real fault in minutes.
Objectives
What you’ll be able to do
- Define metrics: accuracy, latency, cost, safety, security
- Design eval datasets and mixed-method test frameworks
- Run A/B tests and iterate with evidence
- Diagnose prompt failures, hallucinations, model mismatch
- Optimize tokens, latency, and cost–performance trade-offs
- Monitor systems with logging and observability
The story
The Friday deploy
Theo improved a prompt on a Friday. He tried it on three claims — "all three looked better!" — and shipped it. By Tuesday, the water-damage team was in revolt: approvals that used to sail through were being flagged, and nobody could say why. "How do you know the new prompt is better?" Meera asked. "It felt better." She wrote two words on the whiteboard: felt and measured, and drew a wall between them.
They built the inspection line that weekend. Two hundred claims with known-correct outcomes — the golden set, drawn from real adjuster decisions, including the ugly edge cases. Automated checks for format and policy citations. An LLM judge scoring reasoning quality, spot-audited by humans. Every prompt change now walked the line before it touched production, and A/B tests settled arguments that used to be shouting matches.
The next "better" prompt lost to the old one, 61 to 74, on the golden set. It never shipped. Nobody argued. That silence, Meera told Theo, is what evaluation buys: the end of vibes as a deployment strategy.
Metaphor map
Images that stick
| Concept | Image | Why it sticks |
|---|---|---|
| Evaluation | The inspection line | Nothing reaches the loading dock without passing the stations. Not the flagship product, not the "small fix". |
| Golden set | Crash-test dummies | Real crashes, recreated on demand. You don’t wait for a customer to find the failure at highway speed. |
| A/B test | The blind taste test | Two recipes, unlabeled, judged by volume. The chef’s confidence is not an ingredient. |
| Hallucination | The confident tour guide | Fluent, charming, and describing a building that isn’t there. Confidence is not a signal of truth. |
| Latency budget | Traffic signals | Every stage owns a slice of the journey. Find the red light before blaming the road. |
| Cost meter | The fuel gauge | Watched per trip, not per quarter. A leak found at the pump is cheap; found on the highway, it’s a crisis. |
Diagnosis console
Diagnosis console
The exam loves symptom → most-likely-cause questions. Pick a symptom; learn where a working architect looks first.
Retrieval / indexing failure
The only thing that changed is the corpus. A broken re-index, mismatched embeddings, or stale chunks are feeding the model poor context — and a model given wrong context answers wrongly with full confidence.
- Symptom
- Confident but wrong answers began right after a document refresh; latency and model unchanged
- Do first
- the retrieved chunks in a failing trace — are they relevant, current, correctly re-indexed?
- Not this
- model weights, temperature, or the prompt — none of them changed.
Context overflow / eviction
As history grows, early content is truncated, summarized away, or simply drowned. The instruction isn’t being disobeyed — it’s no longer effectively present.
- Symptom
- The assistant ignores instructions that were given early in long conversations
- Do first
- the assembled request near failure: is the instruction still there, and how far from the end?
- Not this
- model quality — the same model obeys the same instruction in short sessions.
Prompt/format brittleness
Intermittent structure failures point to weak format constraints — freeform format instructions instead of enforced structured output, or examples that conflict with the schema.
- Symptom
- Answers are factually fine but the JSON format breaks intermittently under load
- Do first
- failing outputs side-by-side with the format spec; tighten with structured-output enforcement and consistent examples.
- Not this
- retrieval or infrastructure — the content is correct; only the shape wobbles.
Perceived latency (no streaming)
Total generation time may be acceptable, but nothing appears until the full response is done. Streaming and a cached prefix move the first visible token from seconds to near-instant.
- Symptom
- Quality is fine, but users complain the assistant "hangs" before answering
- Do first
- time-to-first-token vs total time in traces; enable streaming, cache the static prefix.
- Not this
- model capability — this is a delivery problem, not an intelligence problem.
Model–task mismatch
The downgraded tier handles surface tasks but lacks depth for multi-step reasoning. Aggregate metrics can hide this — the failures cluster in the hard slice of traffic.
- Symptom
- A cheaper model was swapped in; simple queries are fine but multi-step cases quietly degrade
- Do first
- eval results segmented by difficulty; route complex cases to the capable tier or revert.
- Not this
- prompts or retrieval — they didn’t change; the reasoning budget did.
Optimization lab
Optimization lab
A RAG assistant: 3.5s to first token, 9s total, quality at baseline. Toggle optimizations and watch what each one really buys — and costs. Illustrative numbers; the trade-off shapes are the lesson.
Stable prefix precomputed — cuts TTFT and input cost, zero quality impact
Fewer, better chunks — cheaper and faster, small quality risk
Much cheaper and faster — the only lever that spends quality directly
In production
A product-Q&A bot, measured into shape
In production · retail
A product-Q&A bot, measured into shape
- BASELINE
- 1,000-question golden set from real chats; code checks for price/stock facts, LLM judge for helpfulness, weekly human audit of 50.
- DIAGNOSE
- Post-catalog-refresh accuracy dips traced to stale index chunks in minutes — the trace showed retrieval, not the model, had changed.
- A/B
- New answer format tested on 10% of traffic for two weeks: +6 points helpfulness, no accuracy change, then rolled to 100%.
- OPTIMIZE
- Caching + streaming cut perceived wait dramatically at zero quality cost; a model downgrade was tested, failed the golden set, and was rejected — with evidence.
Storyboard
The film, shot by shot
Every factory has one. The line nothing skips.
Except here. Here, someone tried three examples on a Friday… and felt good.
By Tuesday, the complaints. Fluent answers. Confident answers. Wrong answers.
So we build the dummies. Two hundred real cases, the ugly ones included, answers known in advance.
Now every change walks the line. Facts checked by code. Judgment scored by a judge. The judge audited by humans.
The new prompt scored sixty-one against seventy-four. It never shipped. Nobody argued. That silence is the product.
Flashcards
Recall drill
Select a card to reveal the answer.
Quiz
Exam drill
Question 1
A team wants to replace its prompt with a "clearly better" rewrite. They tried it on four hand-picked examples and all four looked good. What must happen before this ships?
Question 2
You must evaluate an assistant whose answers require judgment (tone, reasoning quality) as well as factual precision (prices, dates). Which evaluation design fits best?
Question 3
An assistant meets its quality bar but misses its cost target by 40%. Traces show a 7k-token static prefix on every call and full responses generated before anything is shown. What should you try FIRST?
Revision
Close your eyes and remember
Felt is not measured. The golden set is scar tissue. Five metrics or it didn't happen. Symptom, then trace, then cause. And the free lunches come first: stream, cache, then — only with proof — touch the model.
Domain 4 · 16% of your exam