Skip to content

Guide

How to compare LLM models

Benchmarks tell you how a model did on someone else's test. The comparison that settles your question is on your own work, and it takes about ten minutes.

Why your own questions beat leaderboards

Public benchmarks measure narrow skills on fixed test sets, and new model versions arrive often. Neither tells you whether a model writes in the tone your team uses, knows your domain, or keeps to the format you need. The only way to find out is to try it on real work.

Five steps

  1. Pick three real tasks: one easy, one typical, one hard. Include the files you'd really attach.

  2. Decide what good means before you look at any answers, and write it down (the checklist below helps).

  3. Run each task on two or three models. In llmwise, ask on one model, then switch models in the same chat and ask again; or use a new chat per model for a blind test.

  4. Score the answers against what you wrote down, not against each other.

  5. Keep the cheapest model that passes. Re-run the three tasks when models change.

What to compare

  • Accuracy

    Is it right? Check the facts and numbers you can check.

  • Following instructions

    Did it do what you asked: the length, the format, the things you said to include or leave out?

  • Editing needed

    How much would you have to fix before you could use it?

  • Tone

    Does it sound the way you need: plain, formal, friendly?

  • Honesty about gaps

    Does it say when it isn't sure, or does it guess confidently?

  • Price

    Once two models are both good enough, the cheaper one wins.

Compare the price, too

In llmwise the model picker shows how many messages you have left on each model, before you send. Start on an everyday model (Claude Haiku 5.5, GPT-6 Luna, DeepSeek V4.1 Flash, and GLM 5.3 Flash) and only move up when it isn't good enough; often it is. See pricing for the allowances.

What you don't have to test

Some differences don't need testing. These come straight from our model catalog.

Every model in llmwise
ModelOn ProOn FreeContext windowImagesPDFsReasoning
Claude Fable 5.1Anthropic31/mo on ProNo1M tokensYesWhole fileYes
Claude Opus 5.5Anthropic62/mo on Pro1 message1M tokensYesWhole fileYes
Claude Sonnet 5.5Anthropic125/mo on ProYes1M tokensYesWhole fileYes
Claude Sonnet 5Anthropic125/mo on ProYes1M tokensYesWhole fileYes
Claude Haiku 5.5Anthropic60/day on ProYes1M tokensYesWhole fileYes
Claude Haiku 4.5Anthropic250/mo on ProYes200K tokensYesWhole fileNo
GPT-6 AstraOpenAI31/mo on ProNo1.05M tokensYesWhole fileYes
GPT-6.1 SolOpenAI125/mo on ProYes1.05M tokensYesWhole fileYes
GPT-6 SolOpenAI125/mo on ProYes1.05M tokensYesWhole fileYes
GPT-6 LunaOpenAI60/day on ProYes1.05M tokensYesWhole fileYes
Gemini 3.1 Pro (preview)Google125/mo on ProYes1.05M tokensYesWhole fileYes
Gemini 3.8 FlashGoogle250/mo on ProYes1.05M tokensYesWhole fileYes
DeepSeek V4.1 FlashDeepSeek60/day on ProYes1.05M tokensYesText onlyYes
DeepSeek V4 ProDeepSeek250/mo on ProYes1.05M tokensNoText onlyYes
Grok 4.7xAI250/mo on ProYes500K tokensYesWhole fileYes
Kimi K3Moonshot125/mo on ProYes1.05M tokensYesText onlyYes
GLM 5.3Z.ai250/mo on ProYes1.05M tokensNoText onlyYes
GLM 5.3 FlashZ.ai60/day on ProYes1.05M tokensYesText onlyYes
Mistral Large 4Mistral250/mo on ProYes1.05M tokensYesText onlyYes
Each badge is how many messages Pro gets on the model: a month’s, or a day’s on an everyday model. Free is a one-time trial of 5 messages on the models marked. “Text only” models get the text of a PDF, not the file.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Questions

Should I trust benchmark leaderboards?

As a rough filter, yes; as a decision, no. A benchmark tells you how a model did on someone else's test. How it does on your questions, in your format, is what you'll live with.

Does the second model see the first one's answer?

In the same chat, yes: after a switch, the next model sees the whole conversation, including the first model's answer. For a blind comparison, ask each model in a new chat of its own.

How many prompts do I need?

Three is enough to start: one easy, one typical and one hard task from your own work. Add more when two models are close.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.