Skip to content

Guide

What is a context window?

A context window is how much a model can take in at once: your messages, its replies, and any files or tool results in the chat, measured in tokens.

Tokens, briefly

Models read text in tokens: short chunks of characters. When llmwise estimates a chat's length it counts about 4 characters of text per token, about 1,600 tokens per image and about 1,500 per PDF page. Each model counts its own way, so treat these as estimates.

What counts toward it

  • Every message in the chat, yours and the model's: the model rereads the conversation each time it answers.

  • Files you attach: the PDF, the image, the CSV.

  • Instructions, such as a persona's.

  • Tool results: search results, pages the model read, the output of code it ran.

When a chat gets long

A model can't read past its window, and long chats are slower and cost more to run, because every reply rereads everything. In llmwise, once the conversation sent to the model passes 64k tokens, each reply counts as 2; past 128k, as 4; and past 200k the chat is full. Your memory and a project's instructions and files count toward a chat's length. The composer says so before you send. Starting a new chat (with a short summary of the old one, if you need it) resets that.

The context window of every model in llmwise

From our model catalog. A PDF counts toward the window whether a model takes the whole file or just its text.

Context windows
ModelMakerContext window
GPT-6 AstraOpenAI1.05M tokens (1,050,000)
GPT-6.1 SolOpenAI1.05M tokens (1,050,000)
GPT-6 SolOpenAI1.05M tokens (1,050,000)
GPT-6 LunaOpenAI1.05M tokens (1,050,000)
Gemini 3.1 Pro (preview)Google1.05M tokens (1,048,576)
Gemini 3.8 FlashGoogle1.05M tokens (1,048,576)
DeepSeek V4.1 FlashDeepSeek1.05M tokens (1,048,576)
DeepSeek V4 ProDeepSeek1.05M tokens (1,048,576)
Kimi K3Moonshot1.05M tokens (1,048,576)
GLM 5.3Z.ai1.05M tokens (1,048,576)
GLM 5.3 FlashZ.ai1.05M tokens (1,048,576)
Mistral Large 4Mistral1.05M tokens (1,048,576)
Claude Fable 5.1Anthropic1M tokens (1,000,000)
Claude Opus 5.5Anthropic1M tokens (1,000,000)
Claude Sonnet 5.5Anthropic1M tokens (1,000,000)
Claude Sonnet 5Anthropic1M tokens (1,000,000)
Claude Haiku 5.5Anthropic1M tokens (1,000,000)
Grok 4.7xAI500K tokens (500,000)
Claude Haiku 4.5Anthropic200K tokens (200,000)

Making the most of it

  • Keep one topic per chat.

  • Attach the pages you need rather than the whole archive.

  • When a chat has done its job, ask for a summary and start fresh with it.

  • Compare how models handle your long documents: the best AI for summarization.

Questions

How much text is a million tokens?

llmwise estimates about 4 characters of text per token, so a million tokens is roughly 4 million characters. Code, numbers and other languages can take more tokens per character.

Why does a long chat count more in llmwise?

Because every reply rereads the whole chat. Past 64k tokens a message does a lot more work, so in llmwise each reply counts as 2 (as 4 past 128k, and a chat stops at 200k), and the composer tells you before you send: “Long chat — each reply counts as 2. Start a new chat to save messages.”

Is a bigger context window always better?

Not always. A model that can take in a million tokens still does its best work on what's relevant. Give it the pages that matter, not everything you have.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.