Est.

Steering AI Agents on an Architectural Model

Prompt wording alone now decides architectural choices that once required deliberate human design.

Staff Writer · · 10 min read
Cover illustration for “Steering AI Agents on an Architectural Model”
Agent Workflows · October 3, 2026 · 10 min read · 2,259 words

A single prompt can now touch dozens of files, pull in a new dependency, scaffold an authentication layer, and wire a retrieval pipeline to a prompt template, all before a human has read a line of what was produced. Researchers behind the paper "Architecture Without Architects: How AI Coding Agents Shape Software Architecture" found that prompt wording alone, with nothing else changed, produced different infrastructure outcomes: a different database, a different framework, a different shape of system, decided in seconds by how a sentence was phrased. That finding deserves attention because it shows architectural choice has moved from a deliberate act performed by a person to a byproduct of language, something teams have started calling vibe architecting.

The scale involved makes this impossible to treat as a rare edge case. Anthropic's 2026 Agentic Coding Trends Report documents sessions that now run dozens of tool calls deep, each one reading a file, writing a block of code, running a command, or iterating on a prior step, across an extended autonomous horizon that a human reviewer never watches in real time. A team that reviews the final diff sees the output of that process. It does not see the dozens of forks in the road where the agent, without anyone's awareness, picked one architecture over another.

The common defense against this observation is that agents simply follow the patterns already present in a codebase, so they aren't really deciding anything novel. The agent produces code that looks like it belongs but quietly breaks the reasoning that justified the original design. That is a more dangerous failure mode than an obviously wrong file, because it passes every surface-level check a reviewer is likely to run.

Architectural drift and the need for an explicit model

Architectural drift is a consequence of how little the agent is allowed to remember, not a sign that the model is insufficiently intelligent. A model can be extraordinarily capable within a session and still drift the moment that session ends, because nothing carries forward except whatever got written into the files themselves.

That absence of memory scales badly. A large production system cannot be compressed that way. "Architecture Without Architects"; frames this precisely as a widening gap: agents build faster than any team can review, and without a persistent design representation sitting alongside the code, there is nothing left to check new work against except the diff in front of a reviewer, which only ever shows what changed, never what the system was supposed to become.

Instruction fidelity, the degree to which a model holds to the constraints it was given across a long task, is a separate axis of quality from raw capability, and a model that drifts from its original constraints mid-task is unsuited to production agentic work no matter how it scores on a benchmark. But fidelity alone does not solve the problem, because even a model with perfect instruction-following can only be as faithful as the specification it is handed. A faithful agent given no architectural model to be faithful to will faithfully reproduce whatever pattern it happens to find first.

Some of this will not improve as models get better. The paper's six prompt-architecture coupling patterns span a range from contingent to fundamental: structured output validation sits on the contingent end and may matter less as models grow more reliable on their own, but tool-call orchestration sits on the fundamental end and persists regardless of how capable any future model becomes, because it concerns how the agent sequences and coordinates actions across a system, not how smart it is within any one of them. The implicit promise that the next model release will quietly fix architectural drift does not hold for that class of pattern. Better prompting narrows the gap for a session. It does not supply the one thing that closes it for good, an explicit record of the architecture that every session can be measured against.

What an explicit architectural model gives agents

An architectural model, in the sense this argument requires, is a structured, machine-readable representation of a system's design that persists across agent sessions and gets consulted the way a human engineer would consult a colleague before making a consequential change. The implication is specific: the limiting factor was never which frontier model a team deployed. It was whether that model had anything reliable to reason against.

The analogy Bain draws holds literally: once written down properly, the explanation only has to happen once, and every subsequent agent session inherits it rather than starting from nothing.

Stripe's internal "minions" system is the clearest public case of this working at scale. That structure produces over a thousand AI-authored pull requests merged per week, every one of them subject to the same human review process applied to any other change at the company. Bain reports Stripe's own conclusion from the arrangement: the specification is the leverage, not the model.

The same logic appears in how codebases get organized, not just in how tasks get specified. More mature versions of this approach treat the agent's system prompt as something closer to a constitutional document, specifying the scope of its authority, the tools it may use, the hard constraints it may not cross, and the conditions under which it must stop and escalate to a person, all layered on top of a structured codebase and a set of steering files.

A shared logic ties these mechanisms together, not a shared tool. Each one takes something that used to live only in a senior engineer's judgment and encodes it into a form an agent can consume directly, and can therefore be held accountable to, at the start of every session rather than reconstructed from scratch inside each one.

What makes verification, not generation, the bottleneck

The constraint on AI-assisted development has moved entirely off code generation and onto the work of checking that generation is correct, and the standard response to that shift, more human eyes on more diffs, is a structurally losing proposition. Bain calls this the shifting bottleneck paradox. The SDLC AI Radar 2026 frames the same dynamic as a coordination failure: organizations that celebrated a spike in raw output, more lines of code, more features shipped, later discovered the bottleneck had simply relocated downstream into integration, QA, and maintenance, surfacing as alignment gaps, duplicated work, and compounding errors wherever AI output wasn't connected to CI/CD pipelines with real validation attached.

The diff is the wrong unit to review, since human defect detection degrades sharply after roughly 400 changed lines, and AI-assisted changes cross that threshold routinely, where unassisted human changes typically stay under it. A reviewer asked to evaluate an AI-generated diff well beyond that threshold is being asked to do a job that research says gets unreliable well before the diff in front of them even ends.

The defects that slip through aren't the same kind a human author tends to produce. Independent benchmarking nonetheless gave it a low completeness score for catching systemic issues, because its analysis is bounded to what changed inside a given pull request, not how that change interacts with the rest of the codebase it was merged into.

The instinct to fix a review bottleneck by adding a second AI reviewer runs into the same wall. A second model reading the same diff inherits the identical structural blind spot as the first: it can describe what changed, but it has no way to check that against what the system's design actually requires, because nothing in the diff tells it what the design requires. Verification built only on diffs, however many reviewers (human or automated) are pointed at them, cannot see architectural intent, because architectural intent does not live in the diff. It lives in the model of the system, and if that model was never written down, there is nothing for any reviewer to check the diff against.

How whole-codebase analysis catches what diff-level review cannot

Closing part of that gap requires review that spans the whole codebase as a connected structure rather than treating a pull request as an isolated artifact. Greptile builds a graph index of an entire codebase specifically so that a review can assess a change's impact beyond the boundaries of the diff itself. In documented hands-on testing, this approach flagged a renamed function signature whose downstream call site the diff under review had never touched, a dependency a diff-only reviewer has no mechanical way of seeing, because the file containing the broken call never appeared in the change set.

This matters well beyond the single example. A caller three files away, a consumer in a separate service, a constraint encoded in a module the agent never opened: none of these are visible from inside the diff, no matter how carefully a human or a model reads it.

A complementary mechanism operates closer to the point where a human reviewer actually intervenes. Revdiff, published by Umputun in April 2026, is a terminal-based diff reviewer that lets a developer annotate specific hunks and lines inside a terminal overlay, then emits a structured JSON payload, file path, hunk, line number, annotation text, an optional tag, that flows directly back to the agent as its next instruction. That structure closes the loop between a human's review and the agent's correction without forcing the agent to re-parse prose comments scattered across a pull request or re-read an entire diff to locate the one line a reviewer actually meant.

What makes the structured payload significant is the same principle running through the architectural model itself: it encodes a reviewer's intent in a form an agent can consume precisely, down to the specific line, rather than leaving that intent to be inferred from a paragraph of free text. Whole-codebase graph analysis and structured review feedback both become more powerful once the graph they operate on is indexed against an explicit design model, because without that model, a tool can tell a team everything that changed across the system but still cannot say whether those changes actually conform to what the design requires. Even at its best, though, this layer of analysis is reactive: it catches a violation after the agent has already committed it to a diff, which sharpens the case for stopping the violation before it reaches a human reviewer.

Architectural constraints belong in CI as deterministic, blocking checks

The only enforcement mechanism that holds up at the speed agents now operate is a CI gate that exits non-zero: a check that runs the same way on every merge, returns the same result given the same input, and blocks a pull request before a human has spent any time reading it. Architecture tests are the concrete form this takes. Tools like ArchUnit on the Java and JVM side, or ESLint configured with custom rules on the JavaScript side, let a team encode a rule like "classes in this package cannot import from that package" directly into the build, so that a violation fails the pipeline automatically rather than waiting to be noticed by a reviewer who happens to know the rule exists.

Practitioners at CodeSai evaluated three strategies for enforcing architectural rules against agent-generated code: architecture tests, review conducted by agents, and documentation paired with skills files. They settled on architecture tests as the primary mechanism, framing the appeal in deterministic terms: a hard-coded rule makes it structurally "difficult to do the wrong thing," rather than merely discouraged. CodeSai had not felt the need to automatically enforce design rules until agents entered the workflow and began taking on delegated tasks; it was the act of delegating to agents, specifically, that made deterministic enforcement necessary in a way it hadn't been when every change passed through a human's hands first.

A CI gate that blocks a failing pull request is a different category of safeguard than a comment left in a review thread. The SDLC AI Radar 2026 names this combination directly as the pattern that works: workflow automation that connects AI output into CI/CD pipelines with validation built in, which is what turns the architectural model from a document someone wrote once into a constraint the system actually enforces on every single change that tries to pass through it.

Spec-driven development restructures the human role from reviewer to architect

The throughline across every mechanism described so far, the steering files, the Stripe-style task specs, the architecture tests, the policy thresholds, is that each one shifts a piece of judgment that used to require a human reading a diff into a form that gets checked before a human ever needs to. That shift changes what a human engineer is actually for. Reading every pull request line by line stops being the leverage point, because the volume of change agents produce has already outpaced what that kind of reading can cover, a fact the review-time data makes plain enough on its own. The spec itself becomes the leverage point: the architectural model a team writes once, the constraints it encodes, the escalation conditions it defines, and the CI gates that hold agents to all three automatically. Stripe's own framing of this, that the specification rather than the model is the leverage, describes a role change as much as a technical one. The engineer's task moves from inspecting what an agent already did toward authoring the model the agent has to work against before it does anything at all, which is a shift from reviewer to architect in the most literal sense available: someone has to design the structure, because the agent, however capable, cannot design it for itself.

Sources

  1. SDLC AI Radar 2026
  2. The Missing Architecture for Agentic Software Development
  3. Architecture Without Architects: How AI Coding Agents Shape Software Architecture
  4. Software development in 2026: A hands-on look at AI agents
  5. Code review best practices for AI-generated diffs
  6. AI Agents for Code Review (2026)
Filed underAgent Workflows

More in Agent Workflows