Engineering
Evaluating prompts like code
Regression suites for instructions: how promptvc catches behavioral drift before production does.
The instruction that drives a production agent is the highest-leverage, least-governed artifact in most AI systems. It is edited in a console by whoever noticed a problem, it is not versioned, it has no tests, and changing it can alter behaviour across every request the system serves.
We would not accept this for a configuration file. We certainly would not accept it for code.
Prompts fail differently
The instinct is to treat a prompt like source and be done. It is nearly right, but prompts break in a way code does not: silently, partially, and in distribution.
A code regression usually announces itself. A prompt regression means that eleven percent of a certain kind of input now gets a subtly worse answer. Nothing throws. Latency is unchanged. The only signal is aggregate quality, which nobody is watching at the resolution required to notice eleven percent.
What a regression suite looks like
Three layers, in increasing cost and decreasing frequency.
Golden cases
A few dozen inputs with known-correct outputs, checked deterministically where possible: does the extraction return the right fields, does the classifier return the right label, is the JSON valid. Fast, cheap, run on every change. These catch gross breakage — a prompt edit that broke the output format — and nothing subtle.
Graded populations
A few hundred representative inputs, scored by a rubric. This is where the eleven percent shows up. The scoring can be model-graded, and model-graded scoring has its own biases, so the rubric must be specific enough that two humans would agree on the score. If your rubric says "is the answer good," you have built a random number generator.
The discipline that made this trustworthy for us was calibration: periodically have humans score a sample the grader has already scored, and track the disagreement rate. When the grader drifts from human judgment, the eval is broken, and you would otherwise never know.
Adversarial and drift sets
Inputs that have broken the system before. Every production incident contributes its triggering input to this set permanently. It is the cheapest institutional memory available and almost nobody maintains it.
Diffing behaviour, not text
The feature that changed how our teams work is behavioural diff. When a prompt changes, the interesting question is not what changed in the instruction — it is which cases changed their answer.
A three-word edit that flips four hundred cases is a large change. A total rewrite that flips none is a refactor. Reviewers should see the second number, and until you build the harness, they cannot.
Version, review, promote
Once the suite exists, the rest is ordinary engineering practice applied to a new artifact. Prompts live in version control. Changes go through review, with the behavioural diff attached. Promotion to production is gated on the suite. A rollback is a revert, not an archaeology expedition.
The objection we hear is that this is heavy for something you want to iterate on quickly. It is the reverse. Iterating quickly is exactly what you cannot do without a way to tell whether you made things better, and teams without a harness do not iterate fast — they iterate nervously, and then stop touching the prompt that works.
Comments
Loading…
Leave a comment