Guide
How to compare LLM models
Benchmarks tell you how a model did on someone else's test. The comparison that settles your question is on your own work, and it takes about ten minutes.
Why your own questions beat leaderboards
Public benchmarks measure narrow skills on fixed test sets, and new model versions arrive often. Neither tells you whether a model writes in the tone your team uses, knows your domain, or keeps to the format you need. The only way to find out is to try it on real work.
Five steps
Pick three real tasks: one easy, one typical, one hard. Include the files you'd really attach.
Decide what good means before you look at any answers, and write it down (the checklist below helps).
Run each task on two or three models. In llmwise, ask on one model, then switch models in the same chat and ask again; or use a new chat per model for a blind test.
Score the answers against what you wrote down, not against each other.
Keep the cheapest model that passes. Re-run the three tasks when models change.
What to compare
Accuracy
Is it right? Check the facts and numbers you can check.
Following instructions
Did it do what you asked: the length, the format, the things you said to include or leave out?
Editing needed
How much would you have to fix before you could use it?
Tone
Does it sound the way you need: plain, formal, friendly?
Honesty about gaps
Does it say when it isn't sure, or does it guess confidently?
Price
Once two models are both good enough, the cheaper one wins.
Compare the price, too
In llmwise the model picker shows how many messages you have left on each model, before you send. Start on an everyday model (Claude Haiku 5.5, GPT-6 Luna, DeepSeek V4.1 Flash, and GLM 5.3 Flash) and only move up when it isn't good enough; often it is. See pricing for the allowances.
What you don't have to test
Some differences don't need testing. These come straight from our model catalog.
| Model | On Pro | On Free | Context window | Images | PDFs | Reasoning |
|---|---|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 31/mo on Pro | No | 1M tokens | Yes | Whole file | Yes |
| Claude Opus 5.5Anthropic | 62/mo on Pro | 1 message | 1M tokens | Yes | Whole file | Yes |
| Claude Sonnet 5.5Anthropic | 125/mo on Pro | Yes | 1M tokens | Yes | Whole file | Yes |
| Claude Sonnet 5Anthropic | 125/mo on Pro | Yes | 1M tokens | Yes | Whole file | Yes |
| Claude Haiku 5.5Anthropic | 60/day on Pro | Yes | 1M tokens | Yes | Whole file | Yes |
| Claude Haiku 4.5Anthropic | 250/mo on Pro | Yes | 200K tokens | Yes | Whole file | No |
| GPT-6 AstraOpenAI | 31/mo on Pro | No | 1.05M tokens | Yes | Whole file | Yes |
| GPT-6.1 SolOpenAI | 125/mo on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes |
| GPT-6 SolOpenAI | 125/mo on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes |
| GPT-6 LunaOpenAI | 60/day on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes |
| Gemini 3.1 Pro (preview)Google | 125/mo on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes |
| Gemini 3.8 FlashGoogle | 250/mo on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes |
| DeepSeek V4.1 FlashDeepSeek | 60/day on Pro | Yes | 1.05M tokens | Yes | Text only | Yes |
| DeepSeek V4 ProDeepSeek | 250/mo on Pro | Yes | 1.05M tokens | No | Text only | Yes |
| Grok 4.7xAI | 250/mo on Pro | Yes | 500K tokens | Yes | Whole file | Yes |
| Kimi K3Moonshot | 125/mo on Pro | Yes | 1.05M tokens | Yes | Text only | Yes |
| GLM 5.3Z.ai | 250/mo on Pro | Yes | 1.05M tokens | No | Text only | Yes |
| GLM 5.3 FlashZ.ai | 60/day on Pro | Yes | 1.05M tokens | Yes | Text only | Yes |
| Mistral Large 4Mistral | 250/mo on Pro | Yes | 1.05M tokens | Yes | Text only | Yes |
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
Questions
Should I trust benchmark leaderboards?
As a rough filter, yes; as a decision, no. A benchmark tells you how a model did on someone else's test. How it does on your questions, in your format, is what you'll live with.
Does the second model see the first one's answer?
In the same chat, yes: after a switch, the next model sees the whole conversation, including the first model's answer. For a blind comparison, ask each model in a new chat of its own.
How many prompts do I need?
Three is enough to start: one easy, one typical and one hard task from your own work. Add more when two models are close.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.