Blocking Merges on Architectural Constraint Violations
AI agents need architectural guardrails at the merge gate, not in code review.

A growing share of production code is now written not by developers typing line by line, but by AI agents that plan and execute multi-file implementation tasks on their own. That shift turns architectural drift into a problem that has to be caught before a merge completes, not one that can wait for a reviewer to notice it after the fact. An agent given a task will return a diff that reads clean and a test suite that passes, while the actual implementation quietly breaks a boundary no test was written to catch.
Picture the task: return a list of users along with their orders. The agent solves it correctly, the output is accurate, and the code ships in minutes. But look at how it got there: one service reaches directly into another service's database rather than going through the boundary that separates them. Nothing about this trips a test. The violation is a relationship between services, not a logic error inside one of them, and relationships between services are exactly the thing code review is worst at catching under time pressure. Worse, a flawed shape like this doesn't stay contained. It gets pulled into a shared library, copied into a scaffold for the next three services, and replicated across a codebase before anyone recognizes that the original pattern was wrong in the first place. An agent optimizes for solving the task in front of it. It has no model of the system's intended shape, and on any codebase large enough to matter, it may not even have that system in its context window to begin with. That mismatch, between what an agent is rewarded for and what a system needs to remain coherent, is the structural reason architectural enforcement has to move to the point where code actually enters the codebase: the merge gate.
Why review-time enforcement fails as the primary architectural defense
Human review has always rested on fragile assumptions: that reviewers remember the relevant conventions, that they apply them consistently, and that there are few enough changes in flight for that discipline to hold. These erode as teams grow and as PR volume rises faster than reviewer bandwidth, and agent output scales faster still. Counterintuitively, the response to more AI-generated code isn't less scrutiny per change. As AI rollout increases total lines of code substantially, developers respond by searching more, not less. The surface each reviewer has to evaluate keeps expanding while the clock they have to do it in does not.
A three-day gap illustrates how this plays out in practice. An agent merges a change cleanly on Monday. The function shape it altered looked self-contained from inside the repository it touched, so the diff passed review without friction. On Wednesday, a downstream team's CI breaks, because the agent had no visibility into the service that depended on the old function shape and no way to know that a boundary it crossed mattered to a system it never saw. That lag is the signature of architectural drift: the damage is real at the moment of merge, but detectable only once something downstream fails.
Review works as governance by inspection: it depends on violations being visible, infrequent, and recognizable to whoever happens to be looking. Agent velocity works against all three conditions at once. Review is not an argument against itself; it is an argument against treating review as the primary mechanism for keeping a system's architecture intact.
What architectural fitness functions are
An architectural fitness function is an automated check that gives an objective, repeatable assessment of whether some property of the system, a boundary, a dependency direction, a layering rule, still holds after a given change. Rather than architecture being something assessed once at design time and then trusted to persist, a fitness function re-evaluates it on every single commit. The shift this represents is from governance by inspection to governance by rule: the constraint is written as executable code, it runs without anyone remembering to invoke it, and it exits non-zero the moment it is violated, with no dependence on a human noticing, recalling, or consistently enforcing a convention.
This matters most precisely where central review doesn't scale. In a microservices architecture or a data mesh, no single team can review every pull request that touches system boundaries, but a fitness function can gate every one of them without a human in the loop at all. The standard being enforced doesn't bend based on who's watching that day or how busy the reviewer queue is.
For teams building with AI agents specifically, the constraint needs to operate at three points, though only one of them is truly load-bearing. Constraints can be embedded directly into the system prompts an agent works from, shaping its behavior before it generates anything. Post-generation validators can run immediately after the agent produces code, failing fast if a constraint was violated. And comprehensive architecture checks can run in CI before anything merges. Of the three, only the CI layer is unconditional and resistant to being bypassed, since prompt-level constraints can be ignored by a model that doesn't attend to them, and post-generation validators can be skipped or misconfigured. CI is where the rule actually holds, regardless of what happened upstream. And it holds the same way for everyone. A fitness function enforces the identical standard whether the code came from a senior engineer, a junior developer working outside their usual domain, or an autonomous agent. The gate has no concept of authorship. It only evaluates the rule.
The concrete tooling: what runs the checks for each major language ecosystem
Every major language ecosystem now has a mature, purpose-built tool for running these checks, so no team can credibly claim that the right tool doesn't exist for their stack. The enforcement principle is identical everywhere. The implementation is ecosystem-specific, and the following is where each ecosystem's option currently stands.
For Java, ArchUnit, released under Apache 2.0 with its latest version, 1.5.0, dated August 4, 2026, lets teams write rules like "no class in the persistence package may depend on a class in the ui package," run those rules as ordinary tests inside CI, and fail the build the moment drift appears. It suits teams that want architectural rules enforced with the same rigor a linter applies to formatting. Because it runs as part of the normal test suite, a PR that introduces drift fails at the moment it is opened rather than during an architecture audit weeks later, by which point the violation has already spread.
For TypeScript and JavaScript, ArchUnitTS has become the leading architecture testing library for TypeScript by download count, checking dependency direction, catching circular dependencies, and enforcing coding standards, with native syntax support for Jest, Vitest, and Jasmine. Alongside it, dependency-cruiser visualizes and enforces the shape of module dependencies, and eslint-plugin-import gates import statements at the linting layer itself. One practitioner recommends running ArchUnitTS and dependency-cruiser together rather than treating them as interchangeable: ArchUnit validates semantic layer constraints while dependency-cruiser enforces module topology, as defense in depth, not alternatives.
For one language ecosystem's tooling, a lint-style checker enforces contracts about which modules are permitted to import which others. For Go, depguard enforces rules about which packages a given package is allowed to depend on.
Across every one of these tools, the underlying pattern never changes: define the constraint as a rule expressed in code, wire that rule into the CI pipeline, and have it exit non-zero the instant it's violated. Code that breaks the rule simply does not merge.
Where in the pipeline the gate belongs
A fitness function only functions as a gate if it sits pre-merge, as a required status check on a protected branch. Anywhere else in the pipeline, it is a report, not a gate.
The Kordi project's CI restructuring shows what it takes to close every gap in that logic. Pull request #1599, merged September 19, 2026, replaced a patchwork of checks with a single stable required check that the merge gate validates against. That check distinguishes a result that is explicitly marked not-applicable from a result that's unexpectedly missing, and it rejects failures, timeouts, cancellations, missing or stale or malformed results, checks run against the wrong revision, and invalid manifests, twenty-four distinct gate fixtures in total, because any one of those gaps left open is effectively an unintended way around the gate. Once the new gate was wired in, the team retired the legacy workflow entirely, a deliberate move to make sure the old, non-blocking path could never quietly become the fallback route when the new one was inconvenient.
EigenScript's CI restructuring, recorded in Issue #1264 with a decision dated September 22, 2026, shows the same principle from the opposite direction: a gate has to be reliable, or teams stop trusting it at all. The project's main branch had gone six days without a clean green run, because one slow lane kept timing out, and once red became the normal state on main, an actual regression had nowhere to stand out. The fix was to separate checks into tiers. Tier 2, the slower lanes covering macOS Intel and various ports, was designed to run post-merge or nightly and open an issue on failure without coloring main red, though subsequent changes moved the macOS suite into the Tier 1 set that gates the merge queue directly. A gate that stays red because a slow or flaky check happens to be required is a gate people learn to route around. The only workable combination is fast, reliable, and mandatory.
GitHub's branch protection ruleset model enforces required checks for every user, including repository owners, by default, and the only way around that enforcement is to be explicitly added to a bypass list, which can grant standing or one-time exemptions. Crucially, the bypass checkbox is a conscious override that the system surfaces, not a silent escape hatch someone can stumble into. That design choice matters: an override that requires a deliberate action leaves a trace and a decision-maker behind it, while a silent exemption leaves none.
A deployment pipeline typically has three natural points where a check could run: pre-merge, post-merge but pre-deploy, and pre-production. Architectural constraints belong specifically at the first of these, because by the time a violation reaches main, the drift has already happened, and every checkpoint after that is cleanup rather than prevention.
The false-positive problem
The strongest objection to merge-blocking architectural enforcement isn't philosophical, it's operational: a gate that blocks legitimate, correct code loses a team's trust faster than it prevents any violation, and once engineers learn that bypassing the gate is the reliable way to get unblocked, the enforcement layer stops functioning in practice no matter what the configuration still says.
A practitioner who built an AI-based architectural review agent for production use found this out directly: false positives turned out to be far more expensive than false negatives, and the mode the team settled on as production-appropriate was the one that achieved zero false positives while accepting a more conservative recall rate. The reasoning is straightforward. Blocking a clean merge because an AI-based detector hallucinated a violation that doesn't exist does more damage to a pipeline's credibility than letting an occasional real violation slip through undetected.
That finding points to a sharp technical distinction. Deterministic, rule-based checks, ArchUnit, import-linter, dependency-cruiser among them, simply do not have a false-positive problem by design, because the rule either matches the code or it doesn't, with no interpretive gray area in between. The risk of false positives belongs specifically to probabilistic or AI-based detection layered on top of those deterministic checks. AI-based review still has a place in the pipeline, but the rules that block a merge outright should be the deterministic ones, with any probabilistic layer advising rather than gating.
The other way a gate loses trust is overreach in the opposite direction: writing a rule for every stylistic preference the team has ever debated turns architectural testing into its own form of overengineering. Rules earn their place by protecting boundaries that matter, categories of mistake that have demonstrably been costly before, and violations that recur often enough to be worth automating, not by codifying preferences a linter already handles. The practical discipline that follows is to start with a small number of high-confidence, high-consequence rules, the ones tied to a violation that has actually broken something in the past, and expand the set only as each existing rule proves itself stable in production.
Why every block must trace to a specific rule
The question every blocked developer asks first is some version of "why did this fail?" The answer has to be a specific rule in a specific file, reconstructable from what's on disk rather than issued as a probabilistic judgment by a model. Every verdict the gate produces has to be reconstructable from what's sitting on disk: the code itself, the text of the rule, and the precise dependency or term that triggered the match. That requirement rules out letting a probabilistic system make the enforcement call on its own. Such systems can retrieve relevant context or recommend a direction, but the decision to block a merge has to trace back to a deterministic rule in a known file producing a known exit code.
Ideally, each rule traces further back still, to an architectural decision record documenting why the boundary exists in the first place, so enforcement isn't tribal knowledge quietly encoded into a check but a documented decision with a machine-readable mechanism attached to it. Without that record, a rule becomes something new developers and new agents simply bump into with no explanation, and the natural response to an unexplained obstacle is to work around it rather than respect it. A rule nobody can explain is a rule that gets deleted the first time it's inconvenient.
Governance guidance for AI systems emerging in 2026 stresses traceability, auditability, and risk management specifically for high-impact systems, including ones that write or modify production code directly, and that regulatory direction happens to point the same way good engineering practice already does. When the architecture of a system lives as a graph in the repository and versions alongside the code it describes, a constraint can reference a specific node in that graph, and the violation message that comes back can point to the exact structural relationship that broke, not a line number in isolation, but a fact about how the system is actually shaped.
How the developer's role shifts under gate-based enforcement
Software development is moving from an activity centered on writing code line by line to one centered on directing agents that write the code, with the engineer accountable for intent, scope, architectural boundaries, and what actually reaches production. That accountability doesn't shrink as agents take on more of the typing. It concentrates on fewer, harder decisions: which boundaries in the system to defend with a blocking rule, which violations deserve an automated check at all, and which exceptions to consciously override rather than silently allow. A gate that enforces architecture on every commit doesn't remove judgment from the process. It frees that judgment to operate where it actually matters, on deciding what the rules should be, rather than on spending every reviewer's finite attention checking whether each individual diff happened to follow them.
Sources
- CI: shared check contract and one blocking merge gate (#1593) by shuyhere · Pull Request #1599 · Kordi-Lab/Kordi
- ci: adopt platform tiers — fast merge-blocking set decides main's color; slow/port lanes post-merge; contributor precheck; unenrolled-test gate · Issue #1264 · InauguralSystems/EigenScript
- Architectural Constraints for AI Agents: Enforcing Structural Patterns in Generated Code