Token Economics & Budgeting
Two API calls can return the same 150-word answer and cost 87 times differently. The price isn't set by the output - it's set by everything sent alongside it: chat history, retrieved documents, system instructions.
Pre-fill vs. payload: Both calls generate the same 150-word response. Call B is 87 times more expensive because the model had to reread 18,000 words of background documents and the entire chat history.
Why output tokens cost 3 to 4 times more than input tokens.
Parallel Processing
When you send a prompt, the GPU reads every input token at once, not one at a time. That parallel processing is why input costs less than output.
Sequential Generation
Output works differently. The model predicts one token, adds it to what it's already written, then predicts the next - one at a time, which costs more per token.
The growing tax of chat history and dynamic retrievals.
System Schemas
Tool definitions and instructions get sent with every request, even though they rarely change. A bloated schema is a tax you pay on every single call.
Chat Accumulation
A multi-turn chat resends the whole conversation every message. If each turn adds 500 tokens, turn 10 is already carrying 5,000 tokens of history.
RAG Payloads
Retrieval injects raw document text into the prompt. One 10,000-token file can multiply the bill 100x for a single question.
Prompt Caching: reducing static input costs by up to 90%.
Modern API providers support prompt caching. If a request shares a long prefix with a previous request (such as a large system prompt, standard codebase conventions, or fixed documentation files), the model retrieves the pre-processed sequence from cache rather than recalculating it.
Uncached Prefix
Every request reads the system prompt and documents from scratch. You pay full price on every turn, even when most of it repeats.
Cached Prefix Match
The repeated system prompt and documents are matched from memory. You pay a small cache-write fee once, then 90% less on every matching call after.
Four strategies to keep context payloads lean.
Sliding Window History
Don't keep every turn forever. Cap it to the last 5-10 messages and drop the rest.
RAG Retrieval Limits
Cap how many chunks or rows a single query can pull in - a few thousand tokens of reference material, not the whole document.
Compact System Schemas
Strip unused fields and shorten descriptions in your tool schemas. Every unnecessary word repeats on every single call.
Hard Output Caps
Set a max token limit on generation. A model stuck in a loop can otherwise write forever on your bill.
"Cost optimization in AI is the discipline of managing context, not just models."
Input tokens are cheap but they add up. Unbounded chat history and documents quietly inflate your bill.
Put static content first. Ordering your prompt so the unchanging parts come before the variable parts gets you more cache hits.