Est.

Context Files for AI Coding Agents in Large Repos

Agents need persistent context files because they can't remember your codebase conventions.

Staff Writer · · 10 min read
Cover illustration for “Context Files for AI Coding Agents in Large Repos”
Agent Workflows · October 6, 2026 · 10 min read · 2,149 words

More than 60,000 repositories have adopted AGENTS.md, a figure that reflects engineering necessity rather than experimentation, as Lulla et al. (2026) document. The context file has stopped functioning as a readme for new hires and started functioning as a dependency the agent cannot work without.

LLM-based agents carry no persistent memory across sessions. Each invocation starts cold, so it has no awareness of the conventions established last week, the mistakes corrected yesterday, or the architectural decisions made last quarter. Someone, or something, has to tell the agent repeatedly and reliably how the project works, because the agent will never simply remember.

Galster et al. (2026) find that context files dominate the configuration landscape for agentic tools and are often the only mechanism a repository uses at all, with AGENTS.md functioning as an interoperable standard that multiple tools now read. That convergence matters because it means the file is no longer a convention specific to one vendor's agent. It has become a shared interface between the codebase and whichever tool is operating on it.

The conceptual change runs alongside the practical one. Documentation written once, for a human reader, just sat beside the code and aged quietly. In 2026, agents are no longer confined to short prompt-and-response exchanges.

What a single context file cannot cover

A single manifest file works fine for a modest codebase. It breaks down once the codebase grows past a certain point, and the way engineering teams have responded to that breakdown says more about the seriousness of the problem than any abstract argument could.

The ceiling is empirical, a function of measurable codebase size. Vasilopoulos (2026) cites a separate empirical study, built on 253 CLAUDE.md files and known as the Agentic Coding Manifests paper, which found that a single file cannot adequately cover a very large codebase. That ratio is striking, but the figure matters less for its size than for what it forces architecturally: a flat file cannot scale linearly with codebase size without becoming a liability of its own, so the structure has to change shape rather than simply grow longer.

The shape that emerges is a hot/cold memory separation, and it is a deliberate design decision rather than an artifact of how the documentation happened to accumulate. Detailed specifications that only matter when the agent is operating in a particular domain stay cold, retrieved on demand rather than loaded by default, which keeps the hot layer from bloating into something an agent has to wade through before it can start work.

Most repositories have not caught up to this structure yet. Galster et al. (2026) confirm that few repositories have adopted advanced mechanisms like subagents and skills, and that most still rely on static context files alone. That means the majority of large repos are carrying a single file further than the evidence suggests it can go.

This level of infrastructure investment looks like an outlier, the product of a researcher working under unusually controlled conditions. But the detail that undercuts that objection is who built it. The person who built Vasilopoulos's 26,000-line context infrastructure came from chemistry, not from software engineering. What drove the scale of the investment was what agent coherence needed to hold a 108,000-line system together, not what an expert engineer's judgment happened to produce.

What the research shows about context files

The research on context files splits cleanly along one line: they measurably change how efficiently an agent works, and the evidence that they change whether an agent completes a task correctly is small at best and contested at worst.

Lulla et al. (2026) find that the presence of AGENTS.md is associated with lower median runtime and reduced output token consumption, while task completion behavior stays comparable with or without the file. That is an efficiency result, not a correctness result, and the distinction carries real weight for anyone deciding where to spend engineering time on these files.

Gloaguen et al. (2026) complicate the picture. The agents follow the file faithfully, into requirements that did not need to be there.

Khatri (2026) adds the most tightly controlled test of the three. The study runs a two-agent ablation across Claude Code and Codex, comparing three injection strategies, no context at all, an always-on full AGENTS.md, and selective retrieval, across 288 evaluated runs. Context strategy does not measurably move correctness on either agent. Khatri's equivalence testing bounds the possible effect at a small number of percentage points for each agent, so this is a precise null result, not a failure to detect anything.

These three findings are not in conflict with one another once the distinction between efficiency and correctness is kept in view. Lulla et al. measured runtime and token use. A task that is trivial for one agent may sit right at the edge of what a different agent can manage, so a study built around a single agent draws its tasks from a different part of that agent's difficulty curve than a study built around another agent, and the two studies land on different conclusions as a result.

Why correctness is unaffected: what agents fail on

Context files leave correctness untouched because the thing agents fail on is a shortfall in implementation skill that no quantity of project documentation can fill in.

Khatri (2026) runs a failure-mode triage on the tasks in the study and finds that agents fail on feature design, pattern selection, and the exact wiring required to connect a new piece of functionality to the rest of the system. They are not facts about how the project is organized, and a context file is built to supply exactly that second category, not the first.

A manipulation probe built into the same study confirms the mechanism directly. What it lacked was the capability to carry the attempt through to a correct implementation, and no additional context supplied before the attempt changed that outcome.

The practical implication follows directly: telling an agent to use the repository pattern for data access, or to follow a particular convention for error handling, does not grant the agent the ability to implement that pattern correctly when it meets a situation the pattern was not written to anticipate. What that same rule cannot do is stop the agent from making a subtly wrong design choice while staying entirely within the boundaries the constraint sets.

Agent capability will likely keep improving, and the gap in implementation skill may well narrow over time. But the underlying point is structural rather than a temporary artifact of where models happen to be today: a context file is a documentation artifact, and documentation does not upgrade capability. Treating the two as interchangeable leads a team to over-invest in writing and polishing the file while under-investing in the verification work that would actually catch the errors the file cannot prevent.

What context files reliably do

Context files earn their keep because they stop specific, repeatable mistakes, the kind that happen when an agent has no memory of the conventions a team settled on in a prior session. That is a narrower claim than making agents smarter, and it is the one the evidence actually supports.

Layered separation, dependency inversion, bounded context boundaries, these are exactly the kind of rule a context file can carry, and Gloaguen et al. confirm behaviorally that agents respect what the file tells them to do: a stated constraint functions as a real constraint in practice rather than a suggestion the agent is free to disregard.

In CI workflows where agents run repeatedly against the same repository, faster and cheaper execution compounds across every run, and the savings accumulate in a way that a single measurement understates.

Vasilopoulos (2026) reports four observational case studies showing how codified context propagates across sessions, preventing failures from recurring and keeping conventions consistent as different sessions touch the same codebase. The mechanism in each case is the guardrail, catching a known failure pattern before it repeats, not the transfer of skill the agent did not already have.

The most actionable principle to come out of this research comes from Gloaguen et al., who find that developer-written context files should describe only minimal requirements. The files that hurt task success are the ones padded with requirements the task did not need; the ones that state precise, narrow constraints do not. A context file earns its place by being short and specific, not by being thorough in the sense of covering every contingency a team can imagine.

ContextCov (2026) points toward where this is heading: a framework built to derive and enforce executable constraints directly from agent instruction files, turning what used to be advisory text into machine-checkable rules. That moves the context file from something an agent merely reads into something a CI pipeline can act on.

Enforcing architectural constraints through CI rather than trusting the context file alone

A rule written into a context file is advisory. A CI check that fails the build when that same rule is violated is enforcement, and only the CI check reliably stops a bad merge from landing.

Static analysis tools already do a version of this today. ArchUnit, built for JVM languages, lets a team specify a rule such as classes in this package cannot import from that package and fail the build automatically when the rule is broken. ESLint, configured with custom rules, covers much of the equivalent ground for JavaScript, handling straightforward import and boundary checks, though matching ArchUnit's full capability, including whole-graph cycle detection and class-level analysis, requires dedicated tools built for that purpose, such as ArchUnitTS or tsarch.

That independence is the point: a gate that only works when the agent behaves as instructed is not really a gate.

The context file and the CI check reinforce each other because they operate at different points in the process. The context file tells the agent what the architecture requires before a single line of code gets written. The CI check catches what the agent got wrong after the code is written. Neither one covers what the other is built to catch, so a team relying on only one of the two is leaving a known gap open.

Even with both layers in place, a reviewer looking at a diff alongside test evidence and a record of the agent's decisions is in a materially better position than a reviewer looking at the diff alone. Research on AI-assisted review backs up the caution this calls for. AI reviewers generate significantly more suggestions than human reviewers do, yet those suggestions get adopted at a lower rate, and more than half of the suggestions that go unadopted turn out to be either factually incorrect or quietly replaced by an alternative the developer wrote instead.

There is a security dimension to build into this from the start: an agent reading pull request content, comments and diffs included, can be targeted by prompt injection hidden inside that content. Agent permissions should be scoped tightly, and any output an agent reviewer produces should be treated as advisory input for a human to weigh, never as a basis for granting that agent write access beyond leaving comments on the PR itself.

Designing a context file that does real work

A context file designed with this research in mind looks different from what most repositories carry today. It is minimal, focused on constraints rather than exposition, tiered according to what the codebase's scale actually demands, and paired with CI enforcement.

Follow Gloaguen et al.'s minimal-requirements principle as a design rule, not as a suggestion. The file should be as short as it can be while still covering the failure patterns a team has already seen repeat.

For repositories that have outgrown the single-file ceiling, the tiered structure Vasilopoulos demonstrates is the model to apply: an always-loaded hot-memory constitution carrying the conventions every session needs, domain-specific agent specifications that embed knowledge directly rather than depending on retrieval to surface it, and a cold-memory collection of detailed specification documents pulled in only when the work calls for them. Architectural invariants and forbidden patterns belong in the hot layer, present on every run. Implementation specifications belong in cold storage, retrieved only when an agent is actually working in that domain. That separation controls how much context bloats a single session while keeping the constraints that matter most in front of the agent at all times.

Wherever possible, pair a constraint stated in the context file with a CI check that enforces it. ContextCov (2026) is a signal of where this is headed: deriving executable constraints directly from the instruction files themselves, so that the file generates its own enforcement.

The context infrastructure has to be versioned with the code it describes, not maintained on a separate schedule. An agent will follow outdated rules with the same confidence it would follow current ones, and that confidence is precisely what makes the file infrastructure in the first place.

Sources

  1. Do Context Files Help Coding Agents?A Two-Agent Ablation Study on Real Repositories
  2. Codified Context: Infrastructure for AI Agents in a Complex Codebase
  3. On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents
  4. Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repositories
  5. Codified Context: Infrastructure for AI Agents in a Complex Codebase
  6. Context Engineering for AI Agents in Open-Source Software
Filed underAgent Workflows

More in Agent Workflows