Engineering
Teaching agents to respect your ADRs
Code generation is easy. Code that honors two hundred prior architecture decisions is not. How we compile ADRs into constraints the Architect agent can’t ignore — and what broke along the way.
An agent that writes code is a solved problem. An agent that writes code your organization would have written is not. The gap between those two sentences is where most enterprise AI pilots quietly die.
We hit it hard in the first months of the AI Factory. The engineer agents were productive — measurably, embarrassingly productive — and roughly a third of what they produced had to be thrown away. Not because it was wrong. Because it contradicted a decision someone had made eighteen months earlier, in a document nobody thought to load into context.
The problem is not memory. It is bindingness.
The obvious fix is to put the decisions in the prompt. We tried that. With two hundred active ADRs across forty services, "put the decisions in the prompt" means either a context window full of governance text that crowds out the actual task, or a retrieval step that surfaces the three most semantically similar decisions — which is not the same as the three that apply.
Semantic similarity is the wrong relation. An ADR about event sourcing in the payments ledger is not similar to a work item about adding a refund endpoint. It is binding on it. Those are different things, and only one of them is a property of the text.
Decisions as typed edges, not documents
What changed things was giving up on ADRs as prose to be retrieved and treating them as nodes with typed relationships to the systems they govern. An ADR in Atlas OS is not a Markdown file that happens to live in a folder. It is an entity with a status, a set of affected components, a supersedes edge, and — the part that matters — a machine-checkable constraint.
When the Planner decomposes a PRD into work items, each item carries the components it touches. The Architect then resolves those components against the graph and pulls every active, non-superseded ADR bound to them. Not the similar ones. The applicable ones.
// constraints are resolved before planning, not retrieved during it
const adrs = await graph.query('adr')
.binding(workItem.systems)
.active();
const plan = architect.plan(workItem, {
constraints: adrs, mode: 'strict'
}); // a violation fails at plan time, not at review
The strict flag is the interesting part. In strict mode a plan that cannot satisfy every bound constraint does not get generated with a warning. It fails, and the failure names the ADR it could not satisfy. The agent's options are to find a compliant approach or to escalate for a decision — which is exactly what we would want a human engineer to do.
What broke
Three things, all of them instructive.
Our ADRs were not decisions
A depressing number of documents in the decision-records folder recorded a discussion and never landed on anything binding. They read as "we considered X and Y, and we are leaning toward X." Leaning is not a constraint. We could not compile them because there was nothing to compile.
Making them machine-readable forced a clarity we should have had anyway. If a decision cannot be expressed as something a system could check, it is probably not finished being made.
Superseding was informal
Engineers had been superseding decisions by writing a newer document and assuming everyone would notice. With agents reading the graph, "everyone would notice" became "nobody notices, forever." We had to backfill supersedes edges across two years of records, and we found four pairs of directly contradictory active decisions that humans had been silently resolving by knowing which one was real.
Strict mode was, at first, unbearable
The first week of strict enforcement, the failure rate was about sixty percent. Most of those failures were correct — the constraints really were being violated — but a meaningful minority came from constraints written more broadly than intended. An ADR that said "all services use the shared audit library" was never meant to apply to the static documentation site.
The fix was scoping, not loosening: narrow the binding edges until the constraint says what its author meant. That work is tedious and it is the actual work. There is no version of this where the ontology does itself.
What it bought
Rework on governed pipelines dropped sharply, but the number we did not expect was review time. Reviewers stopped spending their attention on "this contradicts how we do things" — the pipeline had already caught those — and started spending it on whether the approach was good. That is a much better use of a senior engineer.
The other effect was cultural, and slower. When decisions bind, writing one becomes consequential. People started arguing about ADRs before they were merged rather than after they were ignored.
Which, on reflection, is what the practice was always supposed to produce.
Comments
Loading…
Leave a comment