Glossary · AI Engineering

What is a context window?

Short answer

A context window is the maximum amount of text, measured in tokens, that a large language model can take into account at once: the system prompt, conversation history, supplied documents and its own answer together. Anything beyond the limit has to be shortened, summarised or left out. Current models range from tens of thousands to over a million tokens.

What counts towards the context window

Everything the model sees in one request: system instructions, tool definitions, earlier messages, retrieved documents, the user’s question, and the output it generates. A 200,000-token window holds roughly 150,000 words of English, but long tool definitions and chat histories fill it faster than expected.

Bigger is not automatically better

  • Cost and speed: you pay for every input token on every request, and long prompts are slower.
  • Attention: models can overlook details buried in the middle of very long inputs.
  • Relevance: retrieving the five right passages with RAG usually beats pasting in a whole manual.

Working within the limit

  • Count tokens before sending; the LLM Token Counter estimates size and cost.
  • Summarise older conversation turns instead of resending them in full.
  • Put stable instructions first and use prompt caching where the provider supports it.
  • Ask for concise output when you don’t need long answers.

Published · Updated · By · All terms

Go deeper