Appearance
Spec-Driven Development
What you'll learn
- A one-page markdown spec is the single most valuable artifact you can create before touching an AI coding agent
- Specs work because they constrain the agent's attention: when every decision traces back to a written requirement, the agent stops inventing features you didn't ask for
- The spec template in this lesson takes 15 minutes to fill out and saves hours of correcting an agent that built the wrong thing
- Context pinning with CLAUDE.md keeps project conventions alive across sessions so you don't re-explain your stack every time
The problem
You open Claude Code, describe a feature you want, and the agent starts coding. Ten minutes later you're staring at a PR that adds a dark mode toggle, refactors your auth middleware, and introduces a new dependency you've never heard of. You asked for a search bar. The agent built a search bar plus six things it assumed you'd eventually want, half of which break existing tests.
This is the most common failure mode with AI coding agents: scope drift. Without explicit boundaries, the agent fills in gaps with its own assumptions. It treats your casual description as an invitation to be thorough, and thorough to an LLM means "anticipate every possible future need right now."
The fix is not to write longer prompts. Longer prompts make the problem worse because the agent has more surface area to misinterpret. The fix is a spec: a structured, one-page markdown document that defines exactly what to build, what not to build, and how to verify the result. When the agent's instructions are a spec and not a conversation, scope drift drops dramatically.
Options & when to use each
| Approach | What it's good for | What it costs you | When to pick it |
|---|---|---|---|
| Freeform chat (no spec) | Quick explorations, throwaway scripts, "what does this file do?" | High rework rate. Agent invents requirements, misses constraints, and requires multiple correction rounds for anything nontrivial | 5-minute tasks where the cost of a wrong answer is zero |
| One-page markdown spec | Features that take 30 minutes to 4 hours of agent time. CRUD endpoints, UI components, data pipelines, refactors with clear boundaries | 15 minutes of upfront writing. You need to know what you want before you start | Any feature where getting it wrong costs more than 15 minutes of rework |
| Full PRD / design doc | Multi-week features with multiple stakeholders. Compliance requirements, architectural overhauls, API version bumps | Hours of writing and review. Overkill for solo builders. The spec is longer than the agent needs, and the agent will miss details buried in page 12 | Team environments where a spec already exists. Don't write one for the agent to read |
For solo builders and small teams, the one-page spec is the sweet spot. It's enough structure to keep the agent on track without becoming a writing project in itself.
Build it
Step 1: Create the spec template
Save this as specs/feature-name.md in your project. Replace the bracketed text with your feature's specifics:
markdown
# Feature: [One-line description]
## What this builds
[Three to five sentences describing the finished feature. Describe it from the user's perspective: what they see, what they do, what happens.]
## What this does NOT build
- [Explicitly list anything the agent might assume is in scope but isn't]
- [e.g., "No dark mode," "No admin dashboard," "No email notifications"]
- [If you have a long list here, that's fine. Being explicit about out-of-scope items is the single most important section]
## Technical constraints
- Language: [Python 3.11 / TypeScript / etc.]
- Framework: [FastAPI / Next.js / etc.]
- Database: [PostgreSQL / SQLite / none]
- Packages to use: [list exact packages if it matters]
- Packages to avoid: [list packages the agent tends to reach for but you don't use]
- File structure: [where new files go. e.g., "New endpoints in src/routes/, new models in src/models/"]
## Acceptance criteria
- [ ] [Observable, testable condition. e.g., "GET /api/search?q=term returns JSON array of matching items"]
- [ ] [e.g., "Empty query returns 400 with error message"]
- [ ] [e.g., "Existing tests pass. New test covers the search endpoint"]
- [ ]
## Edge cases to handle
- [What happens when input is empty / too long / malformed]
- [What happens when the resource doesn't exist (404)]
- [What happens when the user isn't authenticated]
- [Race conditions, concurrent requests, rate limiting if applicable]
## How to verify
[Specific commands to run that confirm the feature works. Not "tests pass" but the actual command: `pytest tests/test_search.py -v`]This template forces you to answer the questions the agent would otherwise guess at. The "What this does NOT build" section is the most important: it's the fence that keeps the agent from wandering.
Step 2: Pin project context with CLAUDE.md
Before handing the spec to an agent, make sure your CLAUDE.md tells the agent what it needs to know about your project without being asked. Create or update CLAUDE.md in your project root:
markdown
# Project Context
## Stack
- Python 3.11 with FastAPI
- PostgreSQL via SQLAlchemy 2.0
- Pytest for testing
- Ruff for linting and formatting
## Conventions
- Route handlers go in `src/routes/`. One file per resource.
- Models in `src/models/`. Use SQLAlchemy declarative style.
- Tests in `tests/`. Mirror the src directory structure.
- Never use `print()` for logging. Use `logging.getLogger(__name__)`.
- Always run `ruff check . && ruff format --check .` before committing.
## Commands
- Run server: `uvicorn src.main:app --reload`
- Run tests: `pytest -v`
- Run linter: `ruff check .`
- Format code: `ruff format .`
## Anti-patterns
- Do not introduce new dependencies without asking. The project uses what's in pyproject.toml.
- Do not refactor files outside the scope of the current task.
- Do not generate placeholder code or TODO comments. Build the real thing or don't build it.CLAUDE.md is loaded into every Claude Code session's context. When the agent reads it before processing your spec, it already knows your stack, your conventions, and what not to do. Without CLAUDE.md, the agent makes reasonable guesses that are wrong for your project.
Step 3: Hand the spec to the agent
With the spec written and CLAUDE.md in place, the actual prompt is short:
text
Read specs/payment-webhook.md. Build exactly what the spec describes.
Do not add features not in the spec. Do not refactor files not mentioned in the spec.
If you encounter an ambiguity, stop and ask rather than guessing.If you're using print mode for a one-shot build:
bash
claude -p "Read specs/payment-webhook.md. Build exactly what the spec describes. Do not add features not listed in the spec. Stop and report if anything is ambiguous." \
--allowedTools "Read,Write,Edit,Bash" \
--max-turns 30The key phrase is "Build exactly what the spec describes." This tells the agent that the spec is the complete source of truth, not a starting point for improvisation.
Step 4: Handle ambiguity with a spec review pass
Before building, run a review pass that catches gaps:
text
Read specs/payment-webhook.md. Before writing any code, identify every ambiguous
requirement in the spec. For each ambiguity, state what the spec says, what's unclear,
and list the possible interpretations. Do not build anything.This catches underspecified requirements before the agent commits to the wrong interpretation. If the agent finds three ambiguous clauses, fix the spec and re-run. A spec that passes review with zero ambiguities is ready to build.
Step 5: Verify against the spec, not your memory
After the build, have the agent verify its own work against the spec:
text
Read specs/payment-webhook.md. For each acceptance criterion, confirm whether the
criterion is met. Run the verification commands listed in the spec. Report any
criterion that fails.This closes the loop. The agent built to a spec; now it checks its own work against the same spec. You review the verification report, not the code itself.
What goes wrong
| Mistake | How you notice it | The fix |
|---|---|---|
| Spec is too vague ("add search functionality") | Agent builds something unrelated to what you wanted, or asks 10 clarifying questions before starting | Fill in every section of the template. "What this does NOT build" should have at least three items. If you can't list three out-of-scope things, you haven't thought about the boundaries enough |
| CLAUDE.md is missing or stale | Agent uses the wrong framework, wrong linter, or wrong file conventions | Write CLAUDE.md before you write specs. Update it when you change your stack. The agent only knows what you tell it, and what you tell it once gets loaded every time |
| Agent builds the spec but also builds six extra features | PR is twice the size you expected. Files outside the spec's scope were modified | Add "Do not refactor files not mentioned in the spec" to your prompt. Set --allowedTools "Read,Write,Edit,Bash" without unrestricted access to everything |
| Agent hits a genuine ambiguity and guesses wrong | Built feature handles edge cases differently than you'd expect | Run the spec review pass (Step 4) before building. Fix ambiguities in the spec. A spec that can be interpreted two ways is not done |
| Spec accepts criteria are unverifiable ("the UI looks good") | Agent says it's done but you can't confirm without manual inspection | Every acceptance criterion must be something you can check with a command, a test, or a specific observation: "pytest passes," "curl returns 200," "logs show expected output" |
| Agent runs out of turns on a complex spec | Build stops mid-feature with a partial result | Break the spec into smaller pieces. A spec that takes 40 turns is two specs. Build feature A, verify it, then build feature B. Smaller specs mean smaller blast radius when something goes wrong |
Confirm it worked
Apply the pattern to a real feature in your project:
bash
# 1. Pick a small feature you haven't built yet: a new endpoint, a UI component, a data transformation
# 2. Write the spec using the template (Step 1). Fill in every section.
# 3. Update or create CLAUDE.md with your project conventions (Step 2)
# 4. Run the review pass (Step 4) and fix any ambiguities found
# 5. Hand the spec to Claude Code (Step 3)
# 6. Run the verification pass (Step 5)
# Success: the agent builds exactly what the spec describes, nothing more.
# Failure mode: the agent added features not in the spec. Tighten the "What this does NOT build" section.
# Failure mode: the agent built the wrong thing. Run the review pass again. The spec was still ambiguous.The pattern works when you can hand a spec to an agent, walk away, and come back to a PR that matches your intent without surprises. If you're getting surprises, the spec isn't specific enough yet.
Next: The Alignment Interview -- stress-test your ideas before writing a line of code.