Best AI · Right now
The best AI model right now: all 15, ranked by our tests
In our test runs of all 15 models, DeepSeek V4.1 Flash, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 49 of 50, the most; DeepSeek V4.1 Flash is first on more of the hard ones (20 of 20). The top 6 are within 2 prompts of each other. Here is every model ranked by the same published rule: overall, on the hard prompts alone, and job by job, with what each costs per message.
Test runs and prices checked . Updated .
Short answer
DeepSeek V4.1 Flash, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 49 of 50, the most; DeepSeek V4.1 Flash is first on more of the hard ones (20 of 20). This is our own test, run on September 27, 2026: every prompt, reply and score is published.
Ranked by our test runs
This is our own test: the same prompts sent to every model through llmwise, each reply checked the same way, with every prompt, reply and score published. It ranks the 15 models llmwise offers, run through llmwise, on 50 fixed prompts across 10 jobs. With the top models this close, one prompt can move a model several places: read the scores, not only the places.
| # | Model | Passed | Hard ones | Free trial | On Pro | Cost per reply |
|---|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 FlashDeepSeek | 49 of 50 | 20 of 20 | Yes | 60 a day | $0.0003 |
| 2 | Gemini 3.1 Pro (preview)Google | 49 of 50 | 19 of 20 | Yes | 125 a month | $0.0107 |
| 3 | Claude Opus 5.5Anthropic | 49 of 50 | 19 of 20 | No | 62 a month | $0.0107 |
| 4 | GPT-6 AstraOpenAI | 48 of 50 | 20 of 20 | No | 31 a month | $0.0117 |
| 5 | GLM 5.3Z.ai | 48 of 50 | 19 of 20 | Yes | 250 a month | $0.0007 |
| 6 | GPT-6 LunaOpenAI | 47 of 50 | 20 of 20 | Yes | 60 a day | $0.0001 |
| 7 | Grok 4.7xAI | 47 of 50 | 20 of 20 | Yes | 250 a month | $0.0063 |
| 8 | Claude Sonnet 5Anthropic | 46 of 50 | 19 of 20 | Yes | 125 a month | $0.0041 |
| 9 | Gemini 3.8 FlashGoogle | 46 of 50 | 18 of 20 | Yes | 250 a month | $0.0013 |
| 10 | Kimi K3Moonshot | 46 of 50 | 18 of 20 | Yes | 125 a month | $0.0039 |
| 11 | GPT-6 SolOpenAI | 45 of 50 | 19 of 20 | Yes | 125 a month | $0.0026 |
| 12 | DeepSeek V4 ProDeepSeek | 45 of 50 | 17 of 20 | Yes | 250 a month | $0.0021 |
| 13 | Claude Fable 5.1Anthropic | 45 of 50 | 17 of 20 | No | 31 a month | $0.0200 |
| 14 | Claude Haiku 4.5Anthropic | 43 of 50 | 15 of 20 | Yes | 250 a month | $0.0015 |
| 15 | GLM 5.3 FlashZ.ai | 40 of 50 | 16 of 20 | Yes | 60 a day | $0.0002 |
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
The prompts, every reply and how each was scored: our test runs.
The smartest AI, on our hard prompts
The same 15 models on the hard prompts alone: two a job, from multi-step math to code with edge cases. What each missed shows where it's weakest.
| # | Model | Hard prompts passed | Hard prompts missed |
|---|---|---|---|
| 1 | DeepSeek V4.1 Flash | 20 of 20 | None |
| 2 | GPT-6 Astra | 20 of 20 | None |
| 3 | GPT-6 Luna | 20 of 20 | None |
| 4 | Grok 4.7 | 20 of 20 | None |
| 5 | Gemini 3.1 Pro | 19 of 20 | writing |
| 6 | Claude Opus 5.5 | 19 of 20 | writing |
| 7 | GLM 5.3 | 19 of 20 | writing |
| 8 | Claude Sonnet 5 | 19 of 20 | data analysis |
| 9 | GPT-6 Sol | 19 of 20 | translation |
| 10 | Gemini 3.8 Flash | 18 of 20 | writing and data analysis |
| 11 | Kimi K3 | 18 of 20 | writing and customer support |
| 12 | DeepSeek V4 Pro | 17 of 20 | coding and writing |
| 13 | Claude Fable 5.1 | 17 of 20 | writing and customer support |
| 14 | GLM 5.3 Flash | 16 of 20 | writing, data analysis, and customer support |
| 15 | Claude Haiku 4.5 | 15 of 20 | coding, writing, data analysis, and customer support |
The best AI for each job
Each job's top three by the same rule, with how many of its prompts each passed. The job's own page has every model and every reply.
| Job | Top score | Models with it | First by the rule |
|---|---|---|---|
| Coding | 5 of 5, 2 of 2 hard | 13 of 15 | GLM 5.3 Flash, on the tie-break |
| Writing | 5 of 5, 2 of 2 hard | GPT-6 Luna and GPT-6 Astra | GPT-6 Luna, on the tie-break |
| Math | 5 of 5, 2 of 2 hard | 14 of 15 | GLM 5.3 Flash, on the tie-break |
| Summarization | 5 of 5, 2 of 2 hard | 7 of 15 | DeepSeek V4.1 Flash, on the tie-break |
| Data analysis | 5 of 5, 2 of 2 hard | 11 of 15 | GPT-6 Luna, on the tie-break |
| Customer support | 5 of 5, 2 of 2 hard | 5 of 15 | DeepSeek V4.1 Flash, on the tie-break |
| Translation | 5 of 5, 2 of 2 hard | 14 of 15 | GPT-6 Luna, on the tie-break |
| SQL | 5 of 5, 2 of 2 hard | All 15 | GPT-6 Luna, on the tie-break |
| RAG and answering from documents | 5 of 5, 2 of 2 hard | All 15 | GPT-6 Luna, on the tie-break |
| Agents and tool use | 5 of 5, 2 of 2 hard | 14 of 15 | GPT-6 Luna, on the tie-break |
Every model on every job
Every model, in overall order, on every job: prompts passed out of five.
| Model | Coding | Writing | Math | Summarization | Data analysis | Customer support | Translation | SQL | RAG and answering from documents | Agents and tool use |
|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 5 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Gemini 3.1 Pro | 5 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Claude Opus 5.5 | 5 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| GPT-6 Astra | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 3 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| GLM 5.3 | 5 of 5 | 4 of 5 | 5 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| GPT-6 Luna | 5 of 5 | 5 of 5 | 4 of 5 | 4 of 5 | 5 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Grok 4.7 | 5 of 5 | 3 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Claude Sonnet 5 | 5 of 5 | 4 of 5 | 5 of 5 | 3 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Gemini 3.8 Flash | 5 of 5 | 3 of 5 | 5 of 5 | 5 of 5 | 4 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Kimi K3 | 5 of 5 | 4 of 5 | 5 of 5 | 3 of 5 | 5 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| GPT-6 Sol | 5 of 5 | 4 of 5 | 5 of 5 | 4 of 5 | 5 of 5 | 3 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| DeepSeek V4 Pro | 4 of 5 | 3 of 5 | 5 of 5 | 4 of 5 | 5 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Claude Fable 5.1 | 5 of 5 | 3 of 5 | 5 of 5 | 3 of 5 | 5 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Claude Haiku 4.5 | 3 of 5 | 4 of 5 | 5 of 5 | 5 of 5 | 3 of 5 | 3 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| GLM 5.3 Flash | 5 of 5 | 2 of 5 | 5 of 5 | 3 of 5 | 4 of 5 | 2 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 4 of 5 |
Which AI is best for you
Which is best depends on what you weigh: the results above, read five ways.
Most prompts passed: DeepSeek V4.1 Flash, 49 of 50, first by the rule.
Best of the everyday models: DeepSeek V4.1 Flash (49 of 50), up to 60 messages a day on Pro.
Best of the 12 models in the 5 free messages you get on sign-up: DeepSeek V4.1 Flash (49 of 50).
Lowest running cost of the top five: DeepSeek V4.1 Flash, $0.0003 a reply on average.
Fastest of the top five: GLM 5.3, 1.8 s a reply on average.
One prompt, side by side
A hard data analysis prompt, “Correlation between ad spend and sign-ups”, where the models didn't all pass.
The prompt: Correlation between ad spend and sign-ups (hard)
Here's weekly ad spend and sign-ups as CSV. What's the Pearson correlation coefficient between ad_spend and signups, to 2 decimal places? week,ad_spend,signups 1,500,42 2,800,55 3,650,49 4,1200,71 5,900,60 6,300,35 7,1100,64 8,700,58 Show your working briefly, then end with a line that says "Final answer: " followed by the answer alone.
DeepSeek V4.1 Flash
Passed: Final answer 0.97: right.
**Working** n = 8, Σx = 6150, Σy = 434, Σxy = 357,550, Σx² = 5,352,500, Σy² = 24,516 Numerator: n·Σxy − Σx·Σy = 8(357,550) − 6150(434) = 2,860,400 − 2,669,100 = 191,300 Denominator: - n·Σx² − (Σx)² = 42,820,000 − 37,822,500 = 4,997,500 - n·Σy² − (Σy)² = 196,128 − 188,356 = 7,772 - √(4,997,500 × 7,772) ≈ 197,080.1 r = 191,300 / 197,080.1 ≈ 0.9707 Final answer: 0.97
492 tokens in, 1,515 out (1,302 of them reasoning) · 4.8 s · $0.0013 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Novita ·
Claude Opus 5.5
Passed: Final answer 0.97: right.
**Means** - Mean ad_spend: 6150 / 8 = 768.75 - Mean signups: 434 / 8 = 54.25 **Sums of deviations** - Σ(dx·dy) = 23,912.5 - Σ(dx²) = 624,687.5 - Σ(dy²) = 971.5 **Pearson r** r = 23,912.5 / √(624,687.5 × 971.5) = 23,912.5 / √606,883,906.25 = 23,912.5 / 24,635.01 ≈ 0.9707 This is a strong positive correlation. Final answer: 0.97
607 tokens in, 1,155 out (946 of them reasoning) · 11.5 s · $0.0261 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·
Claude Sonnet 5
Failed: Final answer 0.99; expected 0.97.
**Working:** n=8, mean(x)=768.75, mean(y)=54.25 Computed deviations and cross-products: - Σ(dx·dy) = 24428.125 - Σ(dx²) = 624687.5 - Σ(dy²) = 971.5 r = Σ(dx·dy) / √(Σdx² · Σdy²) = 24428.125 / √(624687.5 × 971.5) = 24428.125 / 24635.01 ≈ 0.9916 **Final answer: 0.99**
605 tokens in, 1,524 out (1,348 of them reasoning) · 13.8 s · $0.0165 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·
How we ranked them
Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand.
Questions
What is the best AI model right now?
In our own test runs on September 2026, DeepSeek V4.1 Flash, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 49 of 50, the most; DeepSeek V4.1 Flash is first on more of the hard ones (20 of 20). The top 6 models passed between 47 and 49, so one prompt can move a model several places: the best one for you depends on your job, which the table by job shows.
What is the smartest AI?
On our hard prompts alone, DeepSeek V4.1 Flash, GPT-6 Astra, GPT-6 Luna, and Grok 4.7 passed 20 of 20 each. "Smartest" here means the hardest prompts we set, from multi-step math to code with edge cases; it's one test, not a verdict on intelligence.
Which AI is best: ChatGPT, Claude or Gemini?
We tested the models, not the apps. Of OpenAI's, GPT-6 Astra passed 48 of 50; of Anthropic's, Claude Opus 5.5 passed 49 of 50; of Google's, Gemini 3.1 Pro passed 49 of 50. ChatGPT, Claude and Gemini add their own tools, limits and settings on top.
How often does this ranking change?
Each time we run the same prompts again, and a model that joins llmwise is ranked once it has been run. The date of the runs is at the top of this page. Runs count for 45 days; after that, this page says they're out of date until we run the prompts again.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.