Spec-Driven Development With AI Agents
Precise specifications give AI agents the guardrails they need to generate reliable code.

Ask an agent to "add a discount code feature to checkout," and watch what happens next: it has to resolve a dozen decisions nobody told it about. Does the discount stack with other promotions already running? Does it apply before or after tax? If someone tries to reuse a spent code, does the system return a 409 or a 422? None of these questions were in the prompt, so the agent guesses, and each guess is a coin flip that compounds against every other guess across the feature. The code that comes out the other end will run. It will often look right in a pull request. Whether it matches what the team meant is a separate question entirely, and that question is where AI-assisted engineering runs into trouble.
Three distinct patterns appear once this kind of work happens at scale. The third is context collapse: across a task that spans multiple sessions or files, the agent contradicts a decision it made earlier in the same project, and the problem gets worse the longer the project runs.
Deepak Babu Piskala's 2026 paper on spec-driven development names the mechanism behind all three: when an AI coding assistant is asked to add a feature from a vague prompt, it has no choice but to guess what the developer wants, and a spec exists to replace that guess with something closer to a contract. The common instinct is to treat this as a prompting problem, something a more careful sentence of instruction would fix. It isn't, because a prompt, however well-written, describes intent exactly once. The limitation is structural, and fixing it requires an artifact with more staying power than a prompt. That artifact is the spec.
Spec as source of truth, code as derived artifact
Spec-driven development, or SDD, does not mean writing more documentation around the same workflow teams already run. It flips who is in charge: the spec governs, the code is generated from it, and when the two disagree, the spec wins and the code gets corrected to match. In conventional development, code quietly becomes the real source of truth. Piskala's paper names this directly as the condition SDD exists to replace.
The inversion is easiest to grasp through a comparison to compiled software: in SDD, the spec plays the role of source code, and the generated code plays the role of the compiled binary, an artifact you regenerate when the spec changes. That comparison also explains why the distinction between a regular requirements document and an SDD spec is not cosmetic. Piskala's paper draws the line sharply: traditional design documents are advisory, developers read them and then write code that hopefully matches, while SDD specs are enforced, meaning tests fail if the code diverges. An SDD spec is written to run as BDD scenarios, API contract tests, or model simulations, and it lives in the repository and in CI, not in a wiki page nobody has opened since the project kicked off.
That shift changes where a developer's time is best spent. Once a spec can be executed against, typing out the implementation stops being the valuable part of the job, because a machine can do that. Defining intent precisely enough that a machine can't get it wrong becomes the actual work. Piskala's paper reports that early adopters of spec-driven workflows saw substantially fewer LLM code-generation errors when the agent worked from a human-refined spec rather than a raw prompt, and teams using structured specs shipped features in far fewer human hours overall, because the agent needed far fewer rounds of correction. The cost of writing the spec is paid once, up front, against savings that appear every time the spec gets used to generate or regenerate code.
Three levels of specification rigor
SDD is not one fixed way of working. Piskala's paper lays out three levels of rigor, spec-first, spec-anchored, and spec-as-source, and which one a team should use depends on how much is riding on the system and how much trust the team can reasonably place in its generation tooling.
Spec-first is the lightest version: the spec kicks off the first round of generation, and after that, the code is free to drift on its own. The overhead is small, and so is the guarantee. Spec-anchored sits a rung higher: the spec and the code evolve together as living documentation, with automated tests enforcing that they stay aligned on an ongoing basis. Piskala's paper identifies this as the sweet spot for most production systems, because it captures most of the benefit of treating the spec as authoritative without requiring blind trust in a fully automated generation pipeline. Spec-as-source is the top level: humans only ever touch the spec, and the code is generated in full and never hand-edited. It removes drift by construction, but it demands generation tooling mature enough to be trusted with that much responsibility, and for most teams, that tooling isn't there yet.
The honest recommendation, and the one the source material backs without hedging, is to aim for spec-anchored. Spec-as-source is where most of the industry's excitement is pointed, but spec-anchored is where the realistic value sits today. Deciding where a given project lands on this ladder comes down to a short set of questions: is this a production system or a prototype, will more than one person or agent be working against the same codebase, does security or compliance or an existing design system constrain what can be built, is the architecture complex enough that getting it wrong would be costly, and is there a maintenance horizon of six months or more ahead of it. When most of those answers point the same direction, spec-anchored is the target, not spec-first and not the aspirational end of the ladder.
What a machine-actionable spec contains
A spec meant for an agent is structured so a human reviewer and an agent extract the same unambiguous meaning from the same document, a writing discipline distinct from narrative prose aimed at a product manager, and different from a product requirements document that merely ran long.
Four things separate a machine-actionable spec from a conventional requirements document. The first is explicit acceptance criteria: every requirement pairs with a concrete, checkable condition, ideally phrased as a Given/When/Then statement or something close enough to a test assertion to be turned into one, rather than a goal like "the feature should work well," which gives an agent nothing to check itself against. The fourth is traceable versioning, keeping the spec in version control next to the code it describes, so a diff to the spec explains why the generated code changed, and a spec that's gone stale is as visible and as fixable as code that's gone stale.
Set against a conversational prompt, a spec differs categorically, not just by degree. A prompt states intent once, in plain language, and trusts the model to remember it. A spec breaks requirements into numbered items, each with its own ID, pairs each one with Given/When/Then acceptance criteria, settles the API contract before anything gets written, lists edge cases by name, and gets edited and regenerated when requirements change instead of being re-explained from scratch in the hope the agent remembers the last conversation.
Consider a webhook delivery feature. None of that level of detail survives in a one-line prompt, and none of it needs to be reinvented by the agent on the fly.
Teams that have settled into this way of working tend to converge on the same seven-section shape for a spec. It opens with a problem statement, one paragraph describing the user-facing issue before any proposed solution. Functional requirements come next, numbered, with each one stated as a single testable statement. A data and API contract section lays out request and response shapes, field types, which fields are required, and every error code, written with the same precision expected of an OpenAPI definition. A final non-goals and exclusions section states what the spec will deliberately not address, closing off the scope the agent might otherwise feel free to expand into.
The SDD workflow: from constitution through implementation, with the gates that stop drift before it compounds
The canonical SDD workflow runs through seven distinct phases, but the discipline that makes it work doesn't come from the phases themselves. It comes from the human review gates sitting between them. Skipping straight from spec to code without reviewing the plan in between lets intent drift survive all the way into a shipped feature, because once code exists, it looks finished, and a finished-looking artifact is much harder to question than a plan on a page.
The golden rule is simple to state and easy to ignore under deadline pressure: never skip from spec straight to code. Catching the same mistake after the code has been written and merged costs hours, sometimes a production incident.
The Clarify phase is where this discipline earns its keep. Its job is to turn an ambiguous brief into a spec an agent can actually implement without filling gaps on its own judgment. Return to the webhook example: Clarify is the step where someone decides whether a failed webhook delivery should be manually retryable from a dashboard (yes, with a Retry button), what should happen to pending events when the destination endpoint gets deleted (cancel the pending retries rather than attempt delivery), and what the timeout on a single delivery attempt should be. Each of those questions, left unanswered, becomes a decision the agent makes arbitrarily and silently. Answered during Clarify, each becomes a line in the spec that the agent and the reviewer both see before a line of code gets written.
A variation on this pattern that's started to appear is "self-spec," where the agent drafts an initial spec from a short, high-level prompt, and a human reviews and refines that draft before any implementation work begins. The value here is the separation it creates between planning and execution: misunderstandings about requirements get caught while they're still sentences on a page, not after they've been compiled into working code. In workflows that split responsibility across multiple specialized agents, one investigating technical feasibility, another handling architectural planning, a third implementing against the task list, a fourth checking the result, the spec is what keeps all of them pointed at the same target across separate sessions and handoffs. The agent topology is a secondary detail. What matters is that the spec is the one document every agent and every human reviewer in the chain is working from.
Verification as the bottleneck when agents write faster than humans can review
A spec, even a good one, run through a disciplined workflow, doesn't eliminate the next problem: agents now produce changes across dozens of files in a single session, and the resulting diffs are often too large for a human reviewer to check carefully.
The scale of this is measurable. Defect detection by human reviewers tends to fall off once a diff passes roughly 400 changed lines, a threshold AI-assisted pull requests cross routinely and unassisted ones typically don't come near. Correctness, security, and whether the change actually fits the system's architecture are much harder to check by eye once a diff reaches that size, regardless of whether the code compiles or parses.
This is exactly where a spec's acceptance criteria and its explicit edge case list stop being planning documents and start being verification instruments. If a spec listed a given edge case, the agent was obligated to handle it, and a test built against that acceptance criterion either passes or proves the agent didn't. The fix for the verification gap is not asking human reviewers to spend more time per pull request, because that doesn't scale against diffs this size. The fix is building the spec so that checking whether generated code matches intent becomes something a machine checks automatically rather than something a person has to hold in their head. Architecture tests wired into a CI/CD pipeline enforce this directly: a rule violation fails the build, and a build failure blocks the change from merging, with no reviewer needing to spot the violation by reading the diff.
Encoding architectural constraints in the spec so generated code cannot silently drift from the design
An acceptance test can confirm that a feature behaves correctly on the surface while the code underneath quietly violates the architecture the team committed to: a service reaching directly into another service's database, a module importing something it was never meant to depend on, a layer boundary an agent didn't know existed because nothing told it. Functional correctness and architectural integrity are different properties, and a spec that only encodes the first leaves the second to chance.
The fix follows the same logic already established for functional requirements: architectural rules need to be written into the spec as checkable constraints, not left as background knowledge the team assumes an agent will infer from the shape of the existing codebase. A rule like "the billing module may not import from the user module directly" is only useful if it's stated as directly as a functional requirement and backed by a test that fails the build the moment it's broken. That test runs the same way a contract test or a BDD scenario runs, inside CI, with no human needing to notice the violation by reading the code. Once those constraints sit in the spec alongside the functional requirements, the two review gates, the one in the workflow and the one in CI, cover both halves of what "correct" means: code that does what the feature was supposed to do, built in a way the architecture was supposed to allow.


