What is a Token?
A token is the basic unit that LLMs read and write. It's roughly 3/4 of a word in English. Every time you use an AI, you're charged by the token - both for what you send and what you get back. Understanding tokens is understanding AI's fundamental economics.
English tokenization is efficient: common words stay whole and punctuation gets its own token. Compound-heavy languages like German split one word into many pieces, driving up token count for the same meaning.
LLMs don't read words. They read tokens.
Tokens are small chunks of text that might be a whole word, part of a word, a punctuation mark, or a space. "Unbelievable" might be 3 tokens: [Un] [believ] [able]. Common words like "the" are 1 token.
A token is roughly 4 characters or 3/4 of a word in English. This is why 100 tokens equals about 75 words. The model never sees the original sentence - it only sees the numbered sequence of token IDs.
You pay per token. Every single one.
Every AI API charges based on tokens: input tokens (your prompt + conversation history + system instructions) and output tokens (the model's response). Output tokens cost more, typically 3 to 4 times more, because generating is harder than reading.
A 10,000-word document you send as context costs money even if the model's answer is one sentence. The input bill runs whether the model uses that context or not.
Token counting is how you understand AI costs. Before you send a prompt, know what's actually inside it.
Everything you send
Your prompt, chat history, system instructions, and any uploaded documents. The model reads these in parallel, so they cost less per token.
Everything the model writes
Generated one at a time, each token depends on all the ones before it. This sequential work costs 3 to 4 times more than reading input.
Models have a token budget. Everything counts against it.
Every model has a context window: the maximum number of tokens it can process in one go. GPT-4 Turbo handles 128K tokens. Claude Opus handles 200K. Gemini and DeepSeek V4 both reach 1 million tokens.
Every message in your conversation, every document you upload, every system instruction - it all counts against this limit. No exceptions, no discounts.
When you hit the limit, older messages get dropped silently or the model refuses to continue. Managing tokens is managing what the model can "see" when it answers you.
Context Windows by Model
Tokenization is not fair across languages.
English is efficient. "Hello, how are you?" takes about 7 tokens. The same phrase in Hindi might take 20 or more tokens. Japanese, Korean, Arabic, and many other languages use 2 to 5 times more tokens than English for the same meaning.
This means non-English users pay more for the same work. A Hindi conversation that costs $0.50 might cost an English user $0.15. The model does the same work either way - the tokenizer just makes some languages more expensive.
It's a known issue with the most common tokenizers, and it means non-English users also hit context limits faster. A document that fits fine in English might overflow the window in Japanese.
Get a feel for the scale before you optimize
Most AI interactions cost fractions of a cent. But RAG setups and long-running conversations can quietly run up token counts without you noticing. Know the rough scale before you start counting pennies.
The rule: watch what goes into context, because you pay for all of it.
"Tokens are the atoms of AI. Everything else - cost, context, capability - is built on top of them."
1 token β 3/4 of an English word. Count tokens, not words, when estimating cost and context usage.
Non-English languages use more tokens. If you work multilingually, factor this into your costs and context planning.