โ† Back to ConceptsVerified: 2026-07-30
Context BasicsAI Basics ยท Question 5

What is a Context Window?

Every LLM has a memory limit. It can only "see" a certain amount of text at once โ€” your prompt, the conversation history, any documents you upload. That limit is the context window. When you exceed it, the model forgets the beginning of the conversation.

CONTEXT WINDOW โ€” THE MODEL CAN ONLY SEE WHAT FITS IN THIS FRAMEMessage 1 (oldest) โ€” DROPPED โ€” outside the windowMessage 2 โ€” DROPPED โ€” outside the windowMessage 3 โ€” Inside window โ€” the model can see thisMessage 4 โ€” Inside windowMessage 5 โ€” Your latest message + System Prompt โ€” inside windowNew messages push old messages out. The model has no memory of anything outside the frame.

Think of the context window as a fixed-width frame sliding over your conversation. What's inside the frame, the model can use. What's outside, it has never seen.

01What Counts Toward the Limit

Everything you send counts โ€” not just your question.

The context window holds: your system prompt (the instructions that set the model's behavior), the entire conversation history (every message you and the model exchanged), any documents or code you uploaded, and your latest message. All of it is converted to tokens and all of it counts against the limit. If you upload a 50,000-token document and ask a 100-token question, the model processes 50,100 tokens โ€” and you pay for all of them.

02Window Sizes Today

Context windows have grown dramatically โ€” but they're still finite.

2022

GPT-3: 2,048 tokens

About 1,500 words. You could send a short article and get a response. Long conversations or documents were impossible.

2023

GPT-4: 8K โ†’ 128K tokens

128K tokens โ‰ˆ 300 pages of text. You could upload entire books. Most production use cases fit comfortably.

2024-2025

Claude: 200K ยท Gemini: 1M

Google pushed the frontier to 1 million tokens โ€” enough to process hours of video transcripts or massive codebases in one pass. Anthropic hit 200K with strong recall throughout the window.

2026

DeepSeek V4: 1M at 34x cheaper

The big shift isn't just bigger windows โ€” it's cheaper ones. Running 1M tokens through DeepSeek V4 costs a fraction of what it costs through GPT-5.5. Large-context workloads are becoming economically viable.

03What Happens When It's Full

When the window fills up, the model makes hard choices.

Most systems handle this by dropping the oldest messages first โ€” a sliding window. Your first messages disappear silently. The model no longer has access to early instructions or context you provided at the start of the conversation. This is why long chat sessions sometimes produce lower-quality responses: the model has lost the setup you gave it at the beginning.

04Practical Limits

Bigger isn't always better. There's a catch.

Models don't use their full context window equally well. Research shows that most models pay strongest attention to the beginning and end of the context โ€” the middle gets less focus. This is called the "lost in the middle" problem. If you put critical information in the middle of a 100K-token document, the model might miss it entirely.

More context also means higher cost and slower responses. Processing 100K tokens costs 50x more than processing 2K tokens and takes longer. The practical rule: use as much context as you need, but no more. Put your most important information at the beginning or end where the model pays the most attention.

05How to Manage It

Three practical rules for working with context windows.

Rule 1

Put Instructions First

System prompts and key instructions go at the start of context โ€” they're least likely to get dropped and the model pays most attention to them.

Rule 2

Summarize Long History

Instead of sending the full 50-message conversation, periodically ask the model to summarize what's been discussed. Send the summary, not the raw log.

Rule 3

Be Selective with Documents

Don't upload entire PDFs when you need one paragraph. Use RAG (retrieval) to find relevant chunks first, then send only those. Your context budget is finite.

"The context window is the model's entire world. Everything outside it doesn't exist."

01

Everything you send counts. System prompts, history, documents, your message โ€” it all consumes the context window and you pay for all of it.

02

Put critical information at the start or end. Models pay less attention to the middle of long contexts โ€” don't bury important details there.