Est.

Multi-Agent Orchestration for Production Codebases

Coordinating parallel agents requires rethinking code review beyond diffs.

Correspondent · · 11 min read
Cover illustration for “Multi-Agent Orchestration for Production Codebases”
Agent Workflows · October 5, 2026 · 11 min read · 2,376 words

Multi-agent systems do not simply generate more code than a single assistant would. They generate interdependent code from several agents at once, and each of those agents works inside its own isolated context window, blind to the shared design the rest of the system depends on. That is the real shift, and it changes what can go wrong in a codebase. A single agent, however fast, still produces one linear stream of changes that a human can trace from start to finish. Multiple agents running at once produce several linear streams that only make sense in relation to each other, and nothing in the process guarantees that relation holds.

The newer coding agents also run for a long time without stopping. They explore a codebase, write changes, run tests, read the results, and adjust, often for minutes or hours before anyone looks at what they produced. A single orchestrated run can touch dozens of files before a human ever opens a pull request. The dominant production pattern compounds the exposure: one coordinator agent dispatches work to specialized sub-agents running in parallel, each sub-agent sees only the slice of the codebase it was assigned, and the coordinator itself tracks whether tasks finished, not whether the finished work fits together architecturally.

The Rakuten deployment on vLLM shows what this looks like at real scale: an agent ran autonomously for seven hours across a massive codebase. No agent involved in that run held the complete design of the system in its context, and no human reviewer could step in afterward and verify structural integrity just by reading what changed. That is not a criticism of the agent or the task; it is a description of the ceiling on what diff-reading can accomplish once a run reaches that length and that scope.

The rest of this piece works through what keeps a codebase coherent when no single agent holds the full architectural model and no human can review output at the speed the agents produce it.

The four orchestration topologies engineers encounter, and their structural risks

How agents are coordinated now matters more for system-level outcomes than which model powers any individual agent. That ordering, topology over model choice, is also where architectural risk concentrates, because the shape of the coordination determines what kind of blind spot gets built into the work.

Four topologies cover most of what engineering teams run into. Sequential, or pipeline, orchestration runs agents one after another, with each agent consuming the previous one's output. The structural risk here is contamination: a design error introduced early moves downstream through every later stage before anyone has a chance to catch it. The failure doesn't stay contained to one part of the system; it spreads through the pipeline, because that is what the pipeline is built to do.

Fan-out, or parallel scatter-gather, orchestration has a coordinator split a task into independent pieces, hand them to agents running at the same time, and then merge the results. The risk here is different: each branch can produce work that is correct on its own terms but incompatible with what another branch is building. No branch can see what the others are doing, so a shared interface or a common abstraction can get modified in two directions at once, and nothing in the fan-out itself notices.

Swarm orchestration, where agents coordinate dynamically without a fixed hierarchy, scales to very large numbers of agents working at once. It also carries the highest structural risk of the four, because no coordinator, not even a partial one, holds a view of the overall design intent. Behavior that emerges from a swarm at scale is, by construction, hard to predict in advance.

Research on this question backs up the claim that topology is a real engineering decision rather than a matter of preference. Work on task-adaptive orchestration, published in 2026 under the name AdaptOrch, treats topology selection as a routing problem: task dependency graphs are analyzed and used to pick the orchestration pattern suited to the task, and topology-aware routing beat static, single-topology baselines by 12 to 23 percent using the same underlying models. The gain came from how the work was structured.

The reconciliation step after a fan-out run has its own limitation. The Adaptive Synthesis Protocol approach to merging parallel agent outputs uses a heuristic consistency score based on embedding similarity, a semantic check that can tell when two outputs disagree in substance. It has no way to tell whether the merged result, even when the pieces agree with each other, violates the architectural constraints of the codebase they were merged into. Semantic agreement and architectural validity are not the same property, and a check built for the first says nothing about the second.

The diff as the wrong unit of architectural review when agents run in parallel

Multi-agent runs produce more code, and that volume is a real problem for review queues, but it is not the central one. A diff shows agent-generated changes as a flat list of line mutations, and that flat list is the wrong shape for understanding what a parallel agent run did to a system's structure. A diff was built for a world where one author made one coherent set of changes for one coherent reason. It was never built to answer what happens when four agents, each blind to the other three, touch the same system at once.

Teams running high-adoption multi-agent workflows report pull requests that have grown dramatically larger and code review times that have stretched out to match. Agent execution is no longer the slow part of the process; reviewing what the agents produced is.

A diff cannot answer the questions that actually matter once agents run in parallel. Did the changes keep the system's layer boundaries intact? Did two agents modify the same shared abstraction in ways that don't agree with each other? Did the branches produced by a fan-out actually compose into something coherent at the architectural level, or just into something that passes its individual tests? None of this is visible in a line-by-line view, because none of it is a line-by-line property. It's a property of the relationships between lines, in different files, written by different agents, that a diff was never designed to represent.

Research on reviewer attention describes the failure this produces. Reviewers cannot direct their attention toward the riskiest parts of a change because the agent gives no signal about its own confidence or the reasoning behind any particular choice it made. Every line looks the same regardless of how much judgment went into it, so the rational response for a reviewer is to read everything with equal care, which is slow and exhausting, and architectural problems that only show up across files the reviewer never compares side by side still go uncaught.

Amazon's experience with this problem is a useful data point. Production outages tied to the volume of AI-generated code led the company to require senior engineer sign-off on AI-assisted changes made by junior and mid-level engineers, along with a 90-day safety reset. That's a human-process fix applied to a tooling gap: the review interface available to the engineers didn't give them a way to see the structural consequences of what the agents had done. Adding another layer of human sign-off doesn't close that gap; it just adds more people looking at the same inadequate representation.

The fix is not reading diffs more carefully. The diff is the wrong abstraction for catching architectural drift once agents are running in parallel, and no amount of review discipline changes what the format can show.

What architectural enforcement requires: intent upstream, deterministic gates downstream

If the diff can't carry the information that matters, the information has to be defined somewhere else, before the agents start, and checked somewhere else, after they finish. Practitioner consensus is converging on a two-layer model: intent and constraints get defined upstream, before any agent generates code, and structural rules get enforced downstream, deterministically, in CI/CD, where a violation blocks the merge.

The upstream layer is about context, not prompting. As generating syntax gets cheap and plentiful, the thing that's actually scarce is architectural control, and that control starts with how intent, constraints, and the system's threat model shape what an agent sees before it writes a single line. The projects that handle this well use constraint files that can be checked mechanically, often in a structured machine-readable format: generated artifacts marked read-only, one single source of truth for the rules, explicit mode matrices for what's allowed where. That's a different thing from a page of prose guidelines: prose guidelines get bypassed or simply forgotten over the course of a long autonomous run, and the drift spreads across many files before anyone notices the damage.

The downstream layer is where enforcement becomes a hard gate rather than a hope. Architecture tests, automated checks that verify the structure of a codebase rather than just its behavior, work as binary pass/fail conditions: a rule violation stops the build, full stop, with no interpretation required. Tools like ArchUnit, or ESLint configured with custom rules, enforce boundaries between layers at build time, and once they're wired into CI/CD, doing the wrong thing becomes mechanically difficult rather than merely discouraged. A check that exits non-zero and blocks a merge is worth more than any amount of review performed after the fact, because it stops drift before it has the chance to compound across the next run.

Neither layer works without the other. Guidelines with no gate behind them drift, because nothing forces an agent, or the human reviewing its output, to actually follow them under time pressure. Gates with no upstream context produce rules that are either too strict to let real work through or too loose to catch anything meaningful, because a gate can only enforce what someone bothered to define clearly first. The two layers depend on each other directly: the upstream layer tells the agent what the system is supposed to look like, and the downstream layer confirms, mechanically, that the output matches.

A living architectural graph as the right enforcement substrate for multi-agent parallelism

Guidelines and architecture tests both run into the same wall when they're checked against static documentation: a document describes what the system looked like when someone last wrote it down, and that stops being true the moment code changes. After a multi-agent run touches forty files in an afternoon, the gap between the documented design and the actual structure of the system is exactly where the risk lives, and a static document has no way to measure a gap it can't see.

What closes that gap is a graph of the codebase as it actually exists: every module, file, class, function, and call relationship, built from the code on disk rather than from someone's notes about the code. A graph like this turns architectural drift into a structural change you can see. It shows which layer boundaries got crossed during a run, which shared abstractions got touched by more than one agent, and whether the branches from a fan-out actually fit together once merged. It answers a different question than a diff does: not what lines changed, but what the structure of the system became.

That only works if the graph is alive. It has to live in the repository and version alongside the code, not sit as a diagram someone drew once in a design doc and never touched again, because a CI check is only as good as the model it checks against, and a stale model will pass code it shouldn't and fail code it shouldn't. The pattern that makes this practical is simple in shape: a command like gr init builds a live architecture graph on demand from the current codebase, and a command like gr check runs structural compliance checks against that graph inside CI. The enforcement layer ends up looking at the exact same graph the developer looked at while planning the work, so there's no gap between the design in someone's head and the constraint the pipeline actually applies.

That shared graph is what makes the two-layer model hold together under parallelism. Developers use the graph upstream, to plan and direct what agents are supposed to build. CI uses the same graph downstream, as the gate that checks what agents actually built. The design and the check are reading from one source, not two, and that single source is what prevents the upstream intent and the downstream enforcement from drifting apart from each other over time.

What developers own in an agent-orchestrated workflow

None of this replaces engineers with autonomous systems. It moves engineers from doing every task by hand to directing, governing, and enforcing design intent across agents that can execute far faster than any person can review line by line. Work that is routine, reversible, and well-scoped is exactly the kind of work that should go to an agent. The architectural model that governs what those agents are allowed to build is not something that can be delegated the same way, and every day a team puts off defining that model explicitly is another day drift accumulates quietly across whatever parallel runs are already underway.

The practical consequence follows directly from that. Teams that put in the work to model their architecture explicitly before running agents, and to enforce that model in CI after, get returns that compound as agent velocity increases: each run gets safer to trust, because the gate around it gets stronger. Teams that skip this step are building technical debt at the same speed their agents are building features, and the debt is harder to see precisely because the features keep shipping.

The orchestration topology a team chooses still matters, and the review process a team runs still matters, but neither one substitutes for an architectural model that is explicit, current, and enforced at the build boundary. A developer who understands what each topology risks, and who keeps a live, enforceable graph of the system's actual structure, is not competing with the agents working alongside them. That developer is the one piece of the system no topology, however well chosen, can supply on its own: the source of the intent the agents are there to carry out.

Sources

  1. AdaptOrch: Task-Adaptive Multi-Agent Orchestration in the Era of LLM Performance Convergence
  2. 2026 Agentic Coding Trends Report How coding agents are reshaping
  3. Effective Strategies for Asynchronous Software Engineering Agents
  4. LLM-Based Multi-Agent Orchestration: A Survey of Frameworks, Communication Protocols, and Emerging Patterns
  5. Rethinking Multi-Agent Collaboration: When More Is Less
  6. Software Engineering in the Agent Era From Trustworthy Change to Human Agent Software Organizations
  7. 3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse
  8. Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems
Filed underAgent Workflows

More in Agent Workflows