โ† Back to ConceptsVerified: 2026-07-30
Token Economics & BudgetingToken Economics: Decoding the API Bill

Token Economics & Budgeting

Two API calls can return the same 150-word answer and cost 87 times differently. The price isn't set by the output - it's set by everything sent alongside it: chat history, retrieved documents, system instructions.

CALL A: STATIC QUERYSystem instructions (100t)User Query (50t)OUTPUT GENERATION (150t)Cost: $0.00075Input: 150 tokens | Output: 150 tokensCALL B: RAG + CHAT CONTEXTSystem instructions (100t)Chat History (3,500t)Retrieved PDF Context (18,000t)User Query (50t)OUTPUT (150t)Cost: $0.06545 (87x increase)Input: 21,650 tokens | Output: 150 tokensVS

Pre-fill vs. payload: Both calls generate the same 150-word response. Call B is 87 times more expensive because the model had to reread 18,000 words of background documents and the entire chat history.

01The Pricing Asymmetry

Why output tokens cost 3 to 4 times more than input tokens.

Input (Pre-fill Phase)

Parallel Processing

When you send a prompt, the GPU reads every input token at once, not one at a time. That parallel processing is why input costs less than output.

Output (Generation Phase)

Sequential Generation

Output works differently. The model predicts one token, adds it to what it's already written, then predicts the next - one at a time, which costs more per token.

02Context Accumulation

The growing tax of chat history and dynamic retrievals.

Layer 1

System Schemas

Tool definitions and instructions get sent with every request, even though they rarely change. A bloated schema is a tax you pay on every single call.

Layer 2

Chat Accumulation

A multi-turn chat resends the whole conversation every message. If each turn adds 500 tokens, turn 10 is already carrying 5,000 tokens of history.

Layer 3

RAG Payloads

Retrieval injects raw document text into the prompt. One 10,000-token file can multiply the bill 100x for a single question.

03The Caching Solution

Prompt Caching: reducing static input costs by up to 90%.

Modern API providers support prompt caching. If a request shares a long prefix with a previous request (such as a large system prompt, standard codebase conventions, or fixed documentation files), the model retrieves the pre-processed sequence from cache rather than recalculating it.

Standard Billing

Uncached Prefix

Every request reads the system prompt and documents from scratch. You pay full price on every turn, even when most of it repeats.

Cost: 10,000 tokens ร— $3.00/M = $0.030 per turn
Optimized Billing

Cached Prefix Match

The repeated system prompt and documents are matched from memory. You pay a small cache-write fee once, then 90% less on every matching call after.

Cost: 10,000 tokens ร— $0.30/M = $0.003 per turn (90% savings)
04Orchestration Rules

Four strategies to keep context payloads lean.

Sliding Window History

Don't keep every turn forever. Cap it to the last 5-10 messages and drop the rest.

RAG Retrieval Limits

Cap how many chunks or rows a single query can pull in - a few thousand tokens of reference material, not the whole document.

Compact System Schemas

Strip unused fields and shorten descriptions in your tool schemas. Every unnecessary word repeats on every single call.

Hard Output Caps

Set a max token limit on generation. A model stuck in a loop can otherwise write forever on your bill.

"Cost optimization in AI is the discipline of managing context, not just models."

01

Input tokens are cheap but they add up. Unbounded chat history and documents quietly inflate your bill.

02

Put static content first. Ordering your prompt so the unchanging parts come before the variable parts gets you more cache hits.