Skip to content

Spec-Driven Development ​

Here is a thing that happens all the time in vibe coding. You type: "Build a user dashboard with filters, sorting, and a CSV export." The agent writes code. It works. You deploy it. Then someone asks: "Does it support filtering by date range?" You check. It does not. You ask: "Does it handle the case where there are zero users?" It crashes with a 500. You ask: "Does the CSV include the email column?" It does not — it includes four random fields and no email at all.

The feature works, but it works for exactly one interpretation of your prompt. And that interpretation — the agent's interpretation — assumed things you never stated and omitted things you assumed were obvious. The result is a feature that passes a demo but fails in production.

The fix is not a longer prompt. The fix is a specification: a document you write before the agent writes any code, that defines exactly what success looks like, what is out of scope, and what edge cases the implementation must survive. When the agent has a spec, it stays on track. When it does not, it wanders.

What you'll learn

  • A spec is not a prompt — it is a contract. It defines what gets built, what does not get built, the behavior other code can depend on, and the edge cases the implementation must handle
  • A good spec makes the agent predictable; a bad spec makes it confidently wrong; no spec makes it invent requirements you never asked for
  • The five-section spec format — what it builds, what it does not build, the contract, edge cases, acceptance criteria — is the minimum that keeps an agent on track across multiple sessions
  • Writing a spec before writing code feels slower, but it prevents the far slower process of discovering missing requirements after the feature is already deployed

The problem: the agent fills in your blanks ​

When you give an agent a vague instruction, it does not ask for clarification. It fills in the blanks with whatever seems most reasonable based on its training data. This is the central challenge of vibe coding: the agent will always produce something, and that something will always look complete, but it may not be what you wanted.

The classic example: "Build a user dashboard." To you, "dashboard" means a table of users with name, email, and signup date, sorted by most recent, with a search bar at the top. To the agent, "dashboard" could mean a metrics page with charts, a list of recent activity, or an admin panel with user management controls. All three are reasonable. Only one is what you wanted.

The agent fills in the blanks in predictable ways:

  • It adds features you never asked for — because the training data associates "dashboard" with charts, it adds a chart library and a set of metrics graphs. You now have a dependency you did not want and code you have to maintain or delete.
  • It skips edge cases you assumed it would handle — "obviously the dashboard should handle empty datasets" is obvious to you, but if you did not say it, the agent does not know it matters. The dashboard crashes on an empty database.
  • It makes architectural decisions without telling you — the agent picks a chart library, a routing pattern, a database query approach, and none of those choices are wrong, but they might be wrong for your project.

The solution is to externalize the requirements. Instead of keeping them in your head and hoping the agent guesses them, write them down in a spec. The spec becomes the single source of truth that the agent checks its work against.

Options & when to use each ​

ApproachWhat it is good forWhat it costs youWhen to pick it
No spec, just a promptTrivial changes: fix a typo, change a color, add a single fieldAgent makes assumptions; feature drifts from your intent; edge cases get missedOne-line fixes and cosmetic changes only
Inline spec in the prompt ("The spec: ...")Quick features where the requirements fit in a paragraphNo persistent record; no way to review the spec independently of the implementationFeatures small enough that the spec and the prompt are the same thing
Separate spec file (this lesson)Features that span multiple prompts, need team review, or have non-obvious edge casesUpfront writing time; spec must be maintained when requirements changeEvery feature beyond a single-prompt fix
Formal spec format (OpenAPI, Gherkin)API contracts or behavior specs that tooling can validate automaticallyHigher overhead; harder to write quickly; overkill for internal featuresPublic APIs or features where automated testing of the spec is valuable

For vibe-coded projects, the separate spec file hits the sweet spot. It is structured enough to keep the agent on track, lightweight enough to write in five minutes, and persistent enough to reference across sessions. A spec written once guides the agent through implementation, review, and testing — three phases where the agent would otherwise drift.

Build it: a spec that actually works ​

Step 1: the five-section format ​

Every spec needs five sections. Not four, not six. Five. Here is why each one matters and what goes in it:

1. What it builds — the feature in concrete terms. Describe what the user sees, what the system does, and what the output looks like. No implementation details (that is the agent's job). No vague adjectives ("clean," "fast," "intuitive"). Just what exists when this feature is done.

2. What it does NOT build — explicit boundaries. This is the most important section for vibe coding because it prevents the agent's most common failure mode: adding features you never asked for. "This spec does not cover user authentication." "This spec does not cover email notifications for new reports." "This spec does not cover mobile responsiveness." Every line in this section is a guardrail.

3. The contract — the technical surface area that other parts of the system can depend on. Route paths, API request and response shapes, database schema changes, new files and directories, configuration additions. The agent uses this section to know where to put things and what they should look like from the outside.

4. Edge cases — scenarios that break a naive implementation. Empty inputs, missing data, concurrent access, very large datasets, authentication failures, rate limiting, browser compatibility. For each edge case, describe the expected behavior. "When the database is empty, the dashboard shows a 'No data yet' message, not a 500 error."

5. Acceptance criteria — specific, testable statements. Not "the dashboard works." Not "users can filter data." Concrete: "A GET request to /api/dashboard?filter=date:2026-08 returns a 200 with a JSON array of metrics for August 2026." "When the database contains zero records, the dashboard renders a page with the text 'No data yet' and a 200 status code." Every acceptance criterion should be something you or the agent can verify with a single command.

Step 2: a real spec, written for the agent ​

Here is a spec for a monthly revenue report feature, written in the five-section format:

markdown
# Monthly Revenue Report

## What it builds

A page at `/reports/revenue` that displays monthly revenue data. The page shows:
- A month selector (dropdown or date picker, user chooses month and year)
- A summary card showing total revenue for the selected month
- A table showing revenue broken down by product category
- A "Download CSV" button that exports the table data

## What it does NOT build

- User authentication or access control — this feature assumes the user is already logged in
- Charts or graphs — the report is tabular data only
- Multi-month comparison — selecting a single month is the only mode
- Email reports or scheduled exports
- Mobile responsiveness beyond what the existing layout already provides
- Historical data import — only data that already exists in the database

## The contract

### Route
- `GET /reports/revenue` — renders the report page
- `GET /api/reports/revenue?month=2026-08` — returns JSON: `{"total_revenue": 12345.67, "categories": [{"name": "Widgets", "revenue": 8000.00}, ...]}`

### Database
- Reads from the existing `orders` and `products` tables
- Joins on `orders.product_id = products.id`
- Aggregates by `products.category`
- No new tables or schema changes

### Files
- `app/routes/reports.py` — new route handlers
- `app/templates/reports/revenue.html` — new template
- `specs/monthly-revenue-report.md` — this spec

### Response format
- API responses return JSON with `total_revenue` (float) and `categories` (array of objects with `name` and `revenue`)
- Errors return `{"error": "message"}` with appropriate HTTP status
- Empty results return `{"total_revenue": 0.0, "categories": []}`

## Edge cases

1. **No data for the selected month** — the API returns zero revenue and empty categories. The page shows "No revenue data for August 2026."
2. **Missing month parameter** — the API defaults to the current month
3. **Invalid month format** — the API returns a 400 with `{"error": "Invalid month format. Use YYYY-MM."}`
4. **Very large dataset** — if a month has more than 10,000 orders, the query should include a LIMIT of 10,000 with a note in the response
5. **CSV with special characters** — category names containing commas or quotes must be properly escaped in the CSV output
6. **Database connection failure** — the API returns a 503 with a generic error message (no database internals exposed)

## Acceptance criteria

1. Visiting `/reports/revenue` shows a month selector and an empty report state
2. Selecting a month with data displays the summary card and category table
3. The CSV download produces a valid CSV file with a header row and the correct data
4. Selecting a month with no data shows the "No revenue data" message with a 200 status
5. `GET /api/reports/revenue?month=2026-08` returns valid JSON matching the contract
6. `GET /api/reports/revenue?month=not-a-date` returns a 400 error
7. All new code follows the conventions in `CLAUDE.md` (snake_case, Pydantic v2, raw SQL, error response format)

This is about 60 lines. It takes ten minutes to write. It covers everything the agent needs to build the feature without guessing.

Step 3: feed the spec to the agent ​

Now the prompt is simple:

Read the spec at @specs/monthly-revenue-report.md and implement it.
Follow the conventions in CLAUDE.md. Write the implementation,
tests that verify every acceptance criterion, and open a PR.

The agent reads the spec, understands the contract, knows the edge cases, and has acceptance criteria it can check its work against. It does not need to guess what "dashboard" means because the spec defines it. It does not add charts because the "What it does NOT build" section explicitly excludes them. It does not crash on an empty database because edge case 1 tells it what to do.

Step 4: the bad spec, for contrast ​

Here is the same feature, written as a bad spec — the kind that causes the agent to wander:

markdown
# Revenue Report

Build a nice revenue report page. It should look clean and professional, with all the data users need. Support date filtering and export.

Make it fast and reliable.

This spec has no contract, no edge cases, no acceptance criteria, and no boundaries. Feed this to the agent and here is what happens:

  • The agent picks a chart library because "all the data users need" sounds like charts
  • The agent adds an authentication system because it is not sure if this page is public
  • The agent invents currency formatting, multi-year comparisons, and a PDF export because "nice" implies feature richness
  • The agent does not handle empty databases because nobody told it to
  • The agent uses a different JSON format than the rest of your API because the contract was never defined

You now have 400 lines of code that do 80% of what you wanted, 20% of things you never asked for, and crash on the edge case your user hits on day one. The ten minutes you saved by not writing a spec will cost you an hour of cleanup.

What goes wrong ​

MistakeHow you notice itThe fix
Writing the spec as a wish list instead of a contractThe agent implements everything literally, and the feature ends up with no coherent structure — each line of the wish list becomes an isolated feature with no relationship to the othersOrganize the spec by contract surface: routes, data shapes, files. The agent uses these to know where things go, not just what to build
Skipping the "What it does NOT build" sectionThe agent adds features you never asked for — charts, auth, notifications, animations — because it is trying to make the feature "complete"Explicitly list what is out of scope. Every line in this section prevents a line of unnecessary code
Writing acceptance criteria that are not testable"The dashboard should be fast" — the agent cannot verify this. "The dashboard should work" — the agent always thinks it worksEvery criterion must be something you can check with a command: "returns a 200," "contains the text 'No data yet'," "CSV has a header row matching the table columns"
Not updating the spec when requirements changeYou realize mid-implementation that you need multi-month comparison, but the spec still says "single month only" — the agent follows the stale specThe spec lives in version control alongside the code. Update it before asking the agent to implement the change, so the spec and the code stay in sync
Making the spec too longYou spend an hour writing a 20-page spec. By the time you finish, you could have built the feature yourselfA spec for a single feature should be readable in under five minutes. If it is longer, the feature is too big — decompose it into smaller features with their own specs
Hiding requirements in the prompt that are not in the specYou tell the agent in the prompt "also add a print-friendly CSS class" but the spec does not mention it. The agent adds it, the next person to read the spec does not know it existsIf it is a requirement, it goes in the spec. The spec is the single source of truth; the prompt is just the delivery mechanism

Confirm it worked ​

Write a spec for a small feature — something you can implement in under an hour. Use the five-section format. Give the spec to the agent with: "Read @specs/feature-name.md and implement it."

After the agent finishes, run through every acceptance criterion. If any criterion fails, the agent did not follow the spec, or the spec was not specific enough. Either way, you have found a problem before it reached production.

Now give the agent the same feature as a one-line prompt with no spec: "Build a feature name page." Compare the output. The no-spec version will almost certainly contain features you did not ask for, miss edge cases the spec version handled, or use patterns that do not match your codebase. This side-by-side comparison is the strongest possible argument for spec-driven development — seeing the same feature built two ways, and one of them is clearly wrong.

Next: Multi-File Project Architecture