Architecture
Ontology migrations without downtime
Evolving the schema of organizational truth while a thousand agents are reading it.
Changing a database schema while traffic flows is a solved problem with a known playbook: expand, migrate, contract. Changing the schema of organizational truth while a thousand agents plan against it is the same problem with one nasty difference — your readers are not just querying the data. They are reasoning about it.
Why the usual playbook is not enough
In an ordinary migration, a reader that does not understand a new column ignores it. Nothing breaks. In an ontology, a reader that does not understand a new relationship type may draw a wrong conclusion rather than no conclusion.
Concretely: we split a single relates-to edge into supersedes, refines, and conflicts-with. During the window when both representations existed, an agent reading only the old edges saw a decision as merely related to another when it was in fact superseded by it. It planned against a retired decision. The graph was internally consistent the entire time. The reasoning on top of it was not.
The pattern that worked
1. Version the ontology, not just the data
Every agent declares the ontology version it understands. The graph knows which version each reader is on. This is the single change that made everything else tractable, and we did it late, which cost us a month.
2. Dual-write, never dual-read
During expansion, writes populate both old and new relationships. Reads go to exactly one representation, chosen by the reader's declared version. Readers that mix representations is precisely how you get the wrong-conclusion failure above.
3. Make the shim explicit and lossy-by-refusal
Old readers see new-world data through a compatibility view. Where the new model draws a distinction the old model cannot express, the view does not collapse it to the nearest old concept — it withholds the edge and flags it. An old agent gets less information, which is safe. It never gets subtly wrong information, which is not.
4. Migrate readers, then contract
Only when every reader has moved to the new version do you drop the old relationships. The graph tells you who is left, because readers declare their version. No guessing, no grep across service configs.
Backfill is a judgment problem
The mechanical part was not the hard part. Splitting one edge type into three means deciding, for every existing edge, which of the three it always meant. Some of those were unambiguous. Several hundred were not, because the original author had used a vague relationship precisely because they were being vague.
We tried to classify them automatically and got about seventy percent confidence, which is a terrible number for something that binds future planning decisions. In the end the ambiguous ones went to the humans who owned the relevant domains, in batches, as a normal work item in the factory. It took three weeks.
The lesson we took: the cost of a vague ontology is paid at migration time, with interest. Every edge whose meaning was obvious in context became an hour of someone's attention once the context was gone.
What we would do differently
Version from day one, even when there is only one version. Refuse vague relationship types at review, the way you would refuse a column called data. And treat the ontology as a published interface with a deprecation policy, because that is what it is — the fact that its consumers are agents rather than services changes nothing except how quietly it fails.
Comments
Loading…
Leave a comment