Appearance
Prompting Patterns That Scale
Most prompting advice is written for one-shot interactions. "Ask the model to explain its reasoning." "Give it a persona." "Use few-shot examples." That advice works when you are generating a single response and throwing it away. It collapses when you are building something that needs to stay coherent across five sessions, three days, and a dozen rounds of iteration.
This lesson is about the patterns that survive scale. System prompts that act as persistent project specs. Context pinning that puts the right information where the agent will see it. Caching strategies that cut cost without cutting correctness. Prompt structures that let an agent pick up where it left off instead of starting from scratch every time.
These are not beginner patterns. These are the patterns that separate a project that survives its second week from one that collapses under its own accumulated confusion.
What you'll learn
- A system prompt is a project spec that loads automatically: put your conventions, constraints, and architecture there, not in every individual prompt
- Context pinning means putting high-value reference material in CLAUDE.md or equivalent so it is always in the agent's attention window
- Prompt caching can cut per-call costs by 50-90%, but only if you structure your prompts so the cached prefix never changes
- Multi-session coherence requires explicit handoff: end every session by asking the agent to write a summary of decisions made and what remains
The problem: prompts that work once and fail twice
You write a prompt on Monday that produces exactly the feature you wanted. On Tuesday you write a second prompt to extend that feature. The agent produces code that conflicts with Monday's output. On Wednesday you try again, and now the agent has forgotten constraints from both previous sessions. By Thursday the codebase is a patchwork of inconsistent decisions that made sense in isolation but do not fit together.
The prompts were fine. The problem is that each prompt operated in a vacuum. The agent had no memory of what it built before, no persistent reference for conventions, and no structural understanding of how the pieces relate. Scaling vibe coding is not about writing better individual prompts. It is about building a system where each prompt inherits the right context from everything that came before it.
Options & when to use each
| Pattern | What it is good for | What it costs you | When to pick it |
|---|---|---|---|
| System prompt as spec | Persistent rules, conventions, and project context loaded automatically every session | Must be maintained; stale specs produce wrong output; too large and it pushes useful context out of the window | Every project. Start with a CLAUDE.md on day one |
| Context pinning | Making sure the agent always sees your most important reference material | Uses context window budget; pinned content that is no longer relevant wastes tokens | Projects with external API docs, database schemas, or compliance rules that every prompt must respect |
| Prefix caching | Cutting per-request cost on repeated context like system prompts and schemas | Adds structural constraint: the cached prefix cannot change between calls | Any project making more than a few dozen API calls per session |
| Session handoff summaries | Letting a new session pick up where the last one left off | Requires discipline to generate and read the summary at the start of each session | Multi-day projects; any project where you cannot finish a feature in one sitting |
| Progressive disclosure | Feeding the agent only the context it needs for the current task, not the whole project | Requires you to know what is relevant before you prompt; risk of omitting something critical | Large projects where the full context would overflow the window |
Build it: the four patterns that keep projects coherent
Pattern 1: The system prompt as a project spec
A system prompt is not a "you are a helpful assistant" preamble. It is a living project specification that the agent reads before every single prompt. Treat it that way.
Put these things in your CLAUDE.md (or equivalent persistent context file):
markdown
# Project: Customer Portal
## Stack
- Backend: Python 3.12, FastAPI, SQLAlchemy 2.0
- Frontend: React 18, TypeScript, Tailwind CSS
- Database: PostgreSQL 16 on localhost:5432
- Package manager: uv (Python), npm (frontend)
## Conventions
- All Python functions must have type hints
- Use Pydantic v2 models for all request/response validation
- Database queries use SQLAlchemy async sessions
- Error responses follow the shape: {"error": "message", "detail": ...}
- No print() statements -- use the logging module
- File naming: snake_case for Python, PascalCase for React components
## Architecture (see ARCHITECTURE.md for full diagram)
- api/ -- FastAPI routes, one file per resource
- models/ -- SQLAlchemy ORM models
- services/ -- business logic, no HTTP concerns
- frontend/src/components/ -- React components
- frontend/src/pages/ -- one component per route
## Current state
- Auth system complete (JWT-based, see api/auth.py)
- User CRUD complete (see api/users.py)
- Currently building: billing integration (branch: feat/billing)The key insight: this is not documentation you write once and forget. It is a spec you update as the project evolves. When you change a convention, update CLAUDE.md. When a new component is complete, add it to the architecture section. When you switch branches, update the current state. The agent's output quality is directly proportional to the accuracy of its persistent context.
Pattern 2: Context pinning
Some reference material needs to be available for every prompt, not every session but every individual prompt. An external API specification. A database schema. A set of compliance rules. You can pin this material in your persistent context so the agent always has it.
The risk is over-pinning. If you pin 20KB of API docs, the agent has 20KB less room for the actual code it is working on. Pin only what every prompt genuinely needs. For reference material that is only relevant to specific tasks, use @file references in individual prompts instead.
Example CLAUDE.md section for pinned context:
markdown
## Pinned reference: Stripe API
The project uses Stripe for payments. Key endpoints:
- POST /v1/payment_intents -- create a payment
- POST /v1/customers -- create a customer record
- Stripe webhooks deliver events to /api/webhooks/stripe
Full API reference pinned at @docs/stripe-api-reference.md.
Use stripe-python SDK, not raw HTTP calls.
Test mode only unless STRIPE_LIVE_MODE=true in .env.Pattern 3: Prefix caching for cost reduction
Both Anthropic and OpenAI cache the prefix of your prompt when it appears across multiple calls. If your system prompt and pinned context total 5,000 tokens and you make 100 calls in a session, you pay full price for those 5,000 tokens once and roughly 10% of the price for the remaining 99 calls. That is a 90% reduction on the cached portion.
The catch is the one rule that breaks caching: the cached prefix must be byte-for-byte identical across calls. Any change, even a single space or a timestamp or a variable that differs between calls, invalidates the cache and you are back to full price.
To preserve caching:
- Put everything static (system instructions, schemas, pinned context) at the very top of each prompt.
- Put everything dynamic (the user's actual question, session IDs, timestamps, the specific code you want edited) at the very bottom.
- Use a consistent prompt template. If you restructure the prompt between calls, the cache breaks.
- Check your provider's billing dashboard for "cache read tokens" or "cached tokens." If that number is zero on repeat calls, something in your prompt prefix is changing.
Anthropic caches prefixes above 1,024 tokens. OpenAI caches above 1,024 tokens as well. Below those thresholds, caching is not worth the structural constraint.
Pattern 4: Session handoff summaries
The most fragile moment in a multi-day vibe coding project is the handoff between sessions. You close your laptop on Tuesday. On Wednesday you open a new session. The agent has no memory of what happened on Tuesday. Unless you give it one.
At the end of every session, ask the agent:
> Write a session summary to SESSION_LOG.md. Include: what we built today,
decisions we made, conventions we established or changed, files created
or modified, and what remains to be built tomorrow.At the start of the next session, the agent reads SESSION_LOG.md (because you reference it in CLAUDE.md or your first prompt). It now has a coherent picture of the project state, the decisions that shaped it, and what it should work on next.
This pattern costs roughly one minute per session and saves roughly one hour of re-establishing context and undoing conflicting decisions.
What goes wrong
| Mistake | How you notice it | The fix |
|---|---|---|
| CLAUDE.md describes a convention you changed last week | The agent follows old rules and produces code that conflicts with the current codebase. You spend time reviewing output that was never going to be right | Update CLAUDE.md when you change conventions. If you changed the error response shape, update the spec. The agent cannot read your mind |
| Pinned context accumulates and crowds out working memory | The agent forgets constraints from early in the session because the context window is mostly pinned reference material | Audit pinned content weekly. If you have not referenced a pinned file in the last three sessions, unpin it and use @file references instead |
| A dynamic value sneaks into the cached prompt prefix | Your per-call token cost stays high even on repeat calls. Cache hit rate in billing dashboard is zero | Move everything dynamic (user query, file paths, timestamps) to the bottom of the prompt. The top must be byte-for-byte identical across calls |
| Skipping the session handoff summary because "I will remember" | You do not remember. The agent starts a new session with no context, makes decisions that conflict with previous work, and you spend the first hour fixing the damage | End every session with a summary. It takes one prompt. The alternative is reconstructing context from memory, which is slower and less accurate |
| Writing a system prompt that is too long | The agent's effective working context shrinks. It loses track of early instructions when processing complex tasks later in the session | Keep system prompts under 2,000 words. Put details in reference files pinned with @file references. The spec is a map, not an encyclopedia |
Confirm it worked
Build a small project over two sessions, one day apart. In session one, scaffold the project, set up CLAUDE.md with conventions, and build one feature. End with a handoff summary. In session two, open a fresh session, load the summary, and build a second feature that depends on the first.
If the second feature integrates cleanly with the first, without the agent inventing new conventions or conflicting with existing code, the patterns work. If the agent produces code that duplicates, conflicts with, or ignores what was built in session one, check your handoff summary: did it capture the decisions and conventions the second session needed to know?
Then check your caching: run the same prompt twice in a row and compare the billing dashboard numbers. If cache hit rate is zero on the second call, audit your prompt template: something in the prefix is changing between calls.
Next: Day 2 -- Building Real Applications -- yesterday you learned the patterns. Tomorrow you build real things with them.