Skip to content
DopeSwagYolo

AI Models

What is a context window in AI?

A context window is the maximum amount of text, counted in tokens, that an AI language model can consider at one time. It includes the prompt, attached files, the conversation so far and the reply the model is writing.

Also known as: context length, token limit

Researched and fact-checked by AI, with no human review. 5 sources listed below. How we verify

Last updated

How it works

Language models process text as tokens. Google's developer documentation says that for its Gemini models one token is about four characters and 100 tokens are roughly 60 to 80 English words. The context window, it says, is the combined limit on input and output tokens.

Anthropic's documentation describes the window as a model's working memory, distinct from the data it was trained on. Everything in a request counts toward it: system instructions, every earlier message, documents, images, tool results and the response being generated. Because each turn is added to the history, a long conversation steadily fills the window.

Anthropic says its API rejects a request whose input alone exceeds the window. It says chat interfaces can instead drop the oldest material first. A beta feature it calls compaction, it says, summarizes earlier parts of a conversation so it can continue.

Why it matters

Window size sets how much material a model can work with in one pass. Google's long-context guide says earlier generative models could process only 8,000 tokens at a time. It says 1 million tokens is roughly 50,000 lines of code or eight average-length English novels.

A bigger window does not guarantee better answers. A 2024 study appeared in Transactions of the Association for Computational Linguistics. It found the models it tested often performed best when relevant information sat at the beginning or end of the input and significantly worse when it sat in the middle. It found this even for models built for long contexts. Anthropic says accuracy and recall degrade as the token count grows, an effect it calls context rot. Google says its models are less accurate when asked to find several specific pieces of information in a long input than when asked for one.

Where things stand in 2026

As of October 2026, Anthropic, Google and OpenAI each list models with context windows of about 1 million tokens. Anthropic's documentation gives 1 million tokens for a group of models that includes Claude Opus 5.5 and Claude Sonnet 5.5. It gives 200,000 for others such as Claude Sonnet 4.5. Google says many Gemini models accept 1 million tokens or more. OpenAI's models page lists a 1.05 million-token context window for each of its three flagship models, GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. These are the companies' own capacity specifications, not independent measures of accuracy at that length.

Sources

  1. Context windows, Anthropic (Claude API documentation)
  2. Understand and count tokens, Google AI for Developers (Gemini API documentation)
  3. Long context, Google AI for Developers (Gemini API documentation)
  4. Lost in the Middle: How Language Models Use Long Contexts, Transactions of the Association for Computational Linguistics, via ACL Anthology
  5. Models | OpenAI API, OpenAI

Articles on AI Models