Architecture

The DAG is the org chart

When work flows between agents, topology is management. What we learned designing the routing layer.

We spent a month designing the routing layer of the AI Factory as an engineering problem. Then we noticed we had been doing management.

The question "which agent picks up this work item, and who reviews it before it moves on" is structurally identical to the question every engineering manager answers when they assign a ticket. We had encoded a management philosophy in a graph and called it orchestration.

Topology is policy

Consider two arrangements. In the first, every engineer agent's output goes directly to the Reviewer, and the Reviewer's approval is the only gate before deployment. In the second, security review runs in parallel with code review, and both must pass.

These are not two implementations of the same pipeline. They are two different organizations. The first believes security is a specialisation of code review. The second believes it is an independent concern with its own authority to block. Whichever you pick, you have made a claim about how your organization handles risk — and you have made it in a routing table rather than in a document anyone will read.

Every edge you draw is a statement about who has authority over what. Draw them carelessly and you will get an organization you did not choose.

Fan-out is delegation. Fan-in is accountability.

The two shapes that recur are worth naming.

Fan-out — one node's output feeding several parallel workers — is delegation. It is fast, and it assumes the work genuinely decomposes. When the Architect fans a work item out to backend, frontend, and platform engineers, it is asserting that those three streams do not need to negotiate with each other. When that assertion is wrong, you get three internally consistent pieces that do not fit together, and the failure surfaces late, at integration, which is the most expensive place for it to surface.

Fan-in — several streams converging on one node — is accountability. Somebody has to hold the whole picture. In our DAG, the Reviewer is a fan-in point, and it is consequently the busiest node in the graph and the one most likely to become a bottleneck. That is not an accident of implementation. Any organization that requires a single coherent judgment about a change will have a bottleneck there, whether it is an agent or a staff engineer.

What we got wrong

Our first topology was too flat. We had eleven agents and very few edges, because edges felt like overhead. The result was an organization where everyone reported to the orchestrator and nobody had context on anyone else's work — the multi-agent equivalent of a startup where the founder is in every meeting.

Throughput looked fine. Quality did not. Agents were making locally reasonable decisions that were globally incoherent, because the topology gave them no way to learn what their peers had decided.

Adding the knowledge graph as a shared substrate fixed more than the routing changes did. Every agent reads and writes the same graph, so the coordination happens through shared state rather than through message passing. That is a familiar answer — it is roughly how a well-run engineering organization works, and roughly why documentation exists.

The uncomfortable implication

If your DAG is your org chart, then redesigning your DAG is a reorganization, and it should be treated with the seriousness of one. We now put topology changes through the RFC process. Someone has to argue for why security should or should not be able to block independently, in writing, before the edge moves.

That sounds like bureaucracy. In practice it has been the cheapest governance we have, because the alternative — discovering your risk posture by reading a routing config six months later — is not cheaper. It is just deferred.

Comments

Loading…

    Leave a comment

    Comments are reviewed before they appear. Your email is never published.