
A Context Window Is the Working Memory of One Request
The context window is the total amount of text a model can hold in mind for a single turn, counted in tokens rather than words. If you paste a long report and the assistant starts losing details, the honest answer is that you have run out of window, and the fix is to trim the input rather than to repeat yourself. Anthropic lists Claude Sonnet 5 with a 1M token window and 128K tokens of output, while Claude Haiku 4.5 lists 200K and 64K, as of Aug 2026.
Everything in this article follows from one fact. That window holds your instructions, the whole conversation so far, any files you attached, and the reply being written, all at the same time.
Tokens Are Not Words and the Gap Is Larger Than People Expect
- ● 1M tokens - about 555,000 words
- ● 200K tokens - about 150,000 words
- ● Code and tables burn more per page
A token is a chunk of text, usually part of a word, and every model counts its window in tokens. Anthropic puts numbers on the conversion and says 1M tokens is roughly 555,000 words on its current tokenizer, while 200K tokens is roughly 150,000 words.
That ratio is a rule of thumb rather than a promise. Code, tables, and languages that do not use the Latin alphabet break into more tokens per visible character, so a spreadsheet dump burns the window faster than plain prose of the same length.
The useful takeaway is a mental exchange rate. Treat a 200K window as a long novel, and treat a 1M window as a short shelf of them.
The Window Is Shared Between What You Send and What You Get Back
Two limits sit side by side on every spec sheet, and people mix them up constantly. The context window is the total, while max output is the ceiling on a single reply.
Anthropic lists Claude Opus 5 at a 1M token window with 128K tokens of maximum output, as of Aug 2026. OpenAI lists GPT-5.6 Sol with a 1.05M token window and 128K tokens of maximum output on its model page, also as of Aug 2026.
So a request cannot fill the entire window with input and still expect a long answer. Room has to be left over, and the tool decides how much when you are not watching.
Why a Long Chat Slowly Forgets How It Started
- ● Every turn re-sends the whole chat
- ● Oldest messages get trimmed first
- ● Cost and latency climb with length
A chat feels continuous to you, but the model gets no memory between turns. Each time you press send, the tool packages the whole conversation again and ships it as one request.
That is why the bill and the latency creep upward in a long session. Turn forty carries thirty nine earlier turns with it, and every one of them occupies the same window your new question needs.
When the total no longer fits, something has to go. Most chat products quietly drop or summarise the oldest messages, which is exactly why the assistant forgets a constraint you set an hour ago while remembering the last thing you typed.
The symptom looks like carelessness and is really arithmetic. Our note on why AI chatbots make things up covers the other half of that behaviour, where a model fills a gap instead of admitting one.
What the Published Numbers Look Like Right Now
These figures come from the official model documentation pages of each provider as of Aug 2026. Confirm current specifications on the official site before you plan around them, since limits change with every model generation.
| Model | Context window | Max output | Rough words in the window | Listed input price |
|---|---|---|---|---|
| Claude Opus 5 | 1M tokens | 128K tokens | About 555,000 | $5 per million tokens |
| Claude Sonnet 5 | 1M tokens | 128K tokens | About 555,000 | $2 per million tokens |
| Claude Haiku 4.5 | 200K tokens | 64K tokens | About 150,000 | $1 per million tokens |
| Claude Fable 5 | 1M tokens | 128K tokens | About 555,000 | $10 per million tokens |
| GPT-5.6 Sol | 1.05M tokens | 128K tokens | Not published in words | Listed on the OpenAI pricing page |
| GPT-5.6 Luna | 1.05M tokens | 128K tokens | Not published in words | Listed on the OpenAI pricing page |
Word estimates come from the Anthropic conversion note and apply to its tokenizer only. Treat them as ranges rather than guarantees, because your own material may pack more tokens per page.
A Big Window Is a Ceiling and Not a Habit
Filling a million token window is possible and rarely wise. Cost scales with the tokens you send, and a conversation that re-sends a large attachment on every turn pays for it on every turn.
Anthropic prices Claude Sonnet 5 at $2 per million input tokens and $10 per million output tokens. It also notes that prompt cache reads cost 10% of the base input price, and that batch requests are 50% off, which is the vendor pointing at the same problem from the billing side.
Quality is the second reason for restraint. A focused prompt with the three relevant pages usually beats the same prompt with sixty pages attached, because the answer no longer depends on the model finding the needle.
The Signs That You Have Run Out of Room
- ● Early instructions stop being followed
- ● Summaries skip the middle
- ● It asks for what you already gave
The window rarely announces itself. It shows up as behaviour that looks like a personality flaw in the tool.
Instructions from early in the chat stop being followed, and the assistant reverts to a default tone or format you corrected an hour ago. A summary covers the beginning and end of a document while skipping the middle, a pattern we look at in reading it yourself versus an AI summary.
Answers about an attached file grow vague, or the tool starts asking for information you already provided. Each of these is the same event wearing different clothes.
What Happens When the Document Will Not Fit At All
Some libraries are simply larger than any window, and stuffing them in is not the answer. The common workaround is retrieval, where a tool searches your documents first and pastes only the passages that look relevant into the prompt.
That design keeps the window small on purpose. It also explains a familiar frustration, since a retrieval tool can miss the passage that mattered and then answer confidently from the passages it did find.
The two approaches suit different jobs. A long window is better when you need the whole document considered together, and retrieval is better when you need one fact out of thousands of pages.
Where an Attached File Actually Goes
Attaching a file in a chat interface usually means the text is extracted and dropped straight into the prompt. A 60 page report can therefore consume a large slice of the window before you have typed your question.
Scanned pages and images behave differently again, since they carry their own token cost once the tool converts them. A folder of screenshots can be heavier than the same information as plain text.
The habit worth building is simple. Attach the pages you need answered rather than the file they live in, and keep the original handy for follow up questions.
Which Window Size Fits Your Work
Everyday questions and short drafts: A 200K token window is far more than enough, and the cheaper, faster model is the better buy. Haiku 4.5 lists 200K tokens, which is roughly 150,000 words.
Long document review and contract reading: Reach for a 1M token model, but paste only the sections in dispute. The window is your ceiling, not your target.
Codebase work across many files: The large window earns its price here, since the model needs several files at once to answer correctly. Expect cost to rise with every added file, not with every question.
Long running research chats: Start a fresh conversation whenever the topic turns. A new chat resets the accumulated history, which is cheaper and sharper than dragging forty turns behind you.
Bulk jobs over many similar inputs: Send them as separate small requests rather than one giant one. Batch pricing exists precisely for this shape of work, and Anthropic lists batch requests at 50% off.
Habits That Stretch a Window Further
Put your instructions at the end of a long prompt as well as the beginning. If earlier material gets trimmed, the rule you care about survives.
Summarise instead of scrolling. When a chat grows long, ask for a short recap of the decisions so far, then start a new conversation with that recap pasted in.
Attach the section rather than the source. One relevant chapter beats a whole manual, and the difference shows up in both the answer and the invoice. If you are choosing between assistants on this basis, our comparison of ChatGPT, Claude and Gemini covers where each one sits.
The Number to Check Before You Blame the Model
Context window, max output, and price per million tokens are three separate lines on a spec sheet, and each answers a different question. The window sets how much the model can consider, the output cap sets how much it can write, and the price sets what a long habit costs.
Read all three before you decide a tool is forgetful. Most complaints about an assistant losing the thread are really a conversation that outgrew its window, and that has a fix you control.
FAQ
What does context window mean in simple terms?
It is the total amount of text a model can hold in mind for one request, counted in tokens. It covers your instructions, the conversation so far, any attached files and the reply being written, all at once.
How many words fit in a 200K token context window?
Anthropic puts 200K tokens at roughly 150,000 words on its current tokenizer, as of Aug 2026. Dense material such as code or tables uses more tokens per page, so treat that as a range rather than a fixed capacity.
Is the context window the same as the maximum response length?
No. The window is the total for input and output together, while max output caps a single reply. Anthropic lists Claude Opus 5 with a 1M token window and 128K tokens of output, which are two separate limits.
Why does a long conversation get slower and more expensive?
The model has no memory between turns, so the tool re-sends the entire conversation with every message. Turn forty carries the previous thirty nine with it, and you pay for those tokens again each time.
Should I always pick the model with the biggest context window?
Not usually. A large window costs more per request and does not improve short tasks, and a focused prompt often beats a bloated one. Pick the big window when the job genuinely needs many files or long documents at once.
Sources
- Anthropic docs: context windows — checked 2026-09-27
- Anthropic pricing — checked 2026-09-27
Some links may be affiliate links. We may earn a commission at no extra cost to you.
This article was written with AI assistance. It is researched and fact-checked, not based on personal hands-on testing unless explicitly stated.
Comments
Post a Comment