Skip to content

Best AI · Right now

The best AI model right now: all 15, ranked by our tests

In our test runs of all 15 models, DeepSeek V4.1 Flash, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 49 of 50, the most; DeepSeek V4.1 Flash is first on more of the hard ones (20 of 20). The top 6 are within 2 prompts of each other. Here is every model ranked by the same published rule: overall, on the hard prompts alone, and job by job, with what each costs per message.

Test runs and prices checked . Updated .

Short answer

DeepSeek V4.1 Flash, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 49 of 50, the most; DeepSeek V4.1 Flash is first on more of the hard ones (20 of 20). This is our own test, run on September 27, 2026: every prompt, reply and score is published.

Ranked by our test runs

This is our own test: the same prompts sent to every model through llmwise, each reply checked the same way, with every prompt, reply and score published. It ranks the 15 models llmwise offers, run through llmwise, on 50 fixed prompts across 10 jobs. With the top models this close, one prompt can move a model several places: read the scores, not only the places.

Models ranked on coding, writing, math, summarization, data analysis, customer support, translation, SQL, RAG and answering from documents, and agents and tool use
#ModelPassedHard onesFree trialOn ProCost per reply
1DeepSeek V4.1 FlashDeepSeek49 of 5020 of 20Yes60 a day$0.0003
2Gemini 3.1 Pro (preview)Google49 of 5019 of 20Yes125 a month$0.0107
3Claude Opus 5.5Anthropic49 of 5019 of 20No62 a month$0.0107
4GPT-6 AstraOpenAI48 of 5020 of 20No31 a month$0.0117
5GLM 5.3Z.ai48 of 5019 of 20Yes250 a month$0.0007
6GPT-6 LunaOpenAI47 of 5020 of 20Yes60 a day$0.0001
7Grok 4.7xAI47 of 5020 of 20Yes250 a month$0.0063
8Claude Sonnet 5Anthropic46 of 5019 of 20Yes125 a month$0.0041
9Gemini 3.8 FlashGoogle46 of 5018 of 20Yes250 a month$0.0013
10Kimi K3Moonshot46 of 5018 of 20Yes125 a month$0.0039
11GPT-6 SolOpenAI45 of 5019 of 20Yes125 a month$0.0026
12DeepSeek V4 ProDeepSeek45 of 5017 of 20Yes250 a month$0.0021
13Claude Fable 5.1Anthropic45 of 5017 of 20No31 a month$0.0200
14Claude Haiku 4.5Anthropic43 of 5015 of 20Yes250 a month$0.0015
15GLM 5.3 FlashZ.ai40 of 5016 of 20Yes60 a day$0.0002
50 prompts per model (coding, writing, math, summarization, data analysis, customer support, translation, SQL, RAG and answering from documents, and agents and tool use), run on September 27, 2026. Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand. Cost per reply is what OpenRouter charged us on average; in llmwise you pay per message.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

The prompts, every reply and how each was scored: our test runs.

The smartest AI, on our hard prompts

The same 15 models on the hard prompts alone: two a job, from multi-step math to code with edge cases. What each missed shows where it's weakest.

Every model on the hard prompts alone
#ModelHard prompts passedHard prompts missed
1DeepSeek V4.1 Flash20 of 20None
2GPT-6 Astra20 of 20None
3GPT-6 Luna20 of 20None
4Grok 4.720 of 20None
5Gemini 3.1 Pro19 of 20writing
6Claude Opus 5.519 of 20writing
7GLM 5.319 of 20writing
8Claude Sonnet 519 of 20data analysis
9GPT-6 Sol19 of 20translation
10Gemini 3.8 Flash18 of 20writing and data analysis
11Kimi K318 of 20writing and customer support
12DeepSeek V4 Pro17 of 20coding and writing
13Claude Fable 5.117 of 20writing and customer support
14GLM 5.3 Flash16 of 20writing, data analysis, and customer support
15Claude Haiku 4.515 of 20coding, writing, data analysis, and customer support

The best AI for each job

Each job's top three by the same rule, with how many of its prompts each passed. The job's own page has every model and every reply.

The top score on each job, and the models that reached it
JobTop scoreModels with itFirst by the rule
Coding5 of 5, 2 of 2 hard13 of 15GLM 5.3 Flash, on the tie-break
Writing5 of 5, 2 of 2 hardGPT-6 Luna and GPT-6 AstraGPT-6 Luna, on the tie-break
Math5 of 5, 2 of 2 hard14 of 15GLM 5.3 Flash, on the tie-break
Summarization5 of 5, 2 of 2 hard7 of 15DeepSeek V4.1 Flash, on the tie-break
Data analysis5 of 5, 2 of 2 hard11 of 15GPT-6 Luna, on the tie-break
Customer support5 of 5, 2 of 2 hard5 of 15DeepSeek V4.1 Flash, on the tie-break
Translation5 of 5, 2 of 2 hard14 of 15GPT-6 Luna, on the tie-break
SQL5 of 5, 2 of 2 hardAll 15GPT-6 Luna, on the tie-break
RAG and answering from documents5 of 5, 2 of 2 hardAll 15GPT-6 Luna, on the tie-break
Agents and tool use5 of 5, 2 of 2 hard14 of 15GPT-6 Luna, on the tie-break
Where several models reach the top score, the rule puts first the smaller message (the everyday models first), then the lower cost per reply: a tie-break on price, not a better answer.

Every model on every job

Every model, in overall order, on every job: prompts passed out of five.

Prompts passed on each job, model by model
ModelCodingWritingMathSummarizationData analysisCustomer supportTranslationSQLRAG and answering from documentsAgents and tool use
DeepSeek V4.1 Flash5 of 54 of 55 of 55 of 55 of 55 of 55 of 55 of 55 of 55 of 5
Gemini 3.1 Pro5 of 54 of 55 of 55 of 55 of 55 of 55 of 55 of 55 of 55 of 5
Claude Opus 5.55 of 54 of 55 of 55 of 55 of 55 of 55 of 55 of 55 of 55 of 5
GPT-6 Astra5 of 55 of 55 of 55 of 55 of 53 of 55 of 55 of 55 of 55 of 5
GLM 5.35 of 54 of 55 of 54 of 55 of 55 of 55 of 55 of 55 of 55 of 5
GPT-6 Luna5 of 55 of 54 of 54 of 55 of 54 of 55 of 55 of 55 of 55 of 5
Grok 4.75 of 53 of 55 of 55 of 55 of 54 of 55 of 55 of 55 of 55 of 5
Claude Sonnet 55 of 54 of 55 of 53 of 54 of 55 of 55 of 55 of 55 of 55 of 5
Gemini 3.8 Flash5 of 53 of 55 of 55 of 54 of 54 of 55 of 55 of 55 of 55 of 5
Kimi K35 of 54 of 55 of 53 of 55 of 54 of 55 of 55 of 55 of 55 of 5
GPT-6 Sol5 of 54 of 55 of 54 of 55 of 53 of 54 of 55 of 55 of 55 of 5
DeepSeek V4 Pro4 of 53 of 55 of 54 of 55 of 54 of 55 of 55 of 55 of 55 of 5
Claude Fable 5.15 of 53 of 55 of 53 of 55 of 54 of 55 of 55 of 55 of 55 of 5
Claude Haiku 4.53 of 54 of 55 of 55 of 53 of 53 of 55 of 55 of 55 of 55 of 5
GLM 5.3 Flash5 of 52 of 55 of 53 of 54 of 52 of 55 of 55 of 55 of 54 of 5

Which AI is best for you

Which is best depends on what you weigh: the results above, read five ways.

  • Most prompts passed: DeepSeek V4.1 Flash, 49 of 50, first by the rule.

  • Best of the everyday models: DeepSeek V4.1 Flash (49 of 50), up to 60 messages a day on Pro.

  • Best of the 12 models in the 5 free messages you get on sign-up: DeepSeek V4.1 Flash (49 of 50).

  • Lowest running cost of the top five: DeepSeek V4.1 Flash, $0.0003 a reply on average.

  • Fastest of the top five: GLM 5.3, 1.8 s a reply on average.

One prompt, side by side

A hard data analysis prompt, “Correlation between ad spend and sign-ups”, where the models didn't all pass.

The prompt: Correlation between ad spend and sign-ups (hard)
Here's weekly ad spend and sign-ups as CSV. What's the Pearson correlation coefficient between ad_spend and signups, to 2 decimal places?

week,ad_spend,signups
1,500,42
2,800,55
3,650,49
4,1200,71
5,900,60
6,300,35
7,1100,64
8,700,58

Show your working briefly, then end with a line that says "Final answer: " followed by the answer alone.
  • DeepSeek V4.1 Flash

    Passed: Final answer 0.97: right.

    **Working**
    
    n = 8, Σx = 6150, Σy = 434, Σxy = 357,550, Σx² = 5,352,500, Σy² = 24,516
    
    Numerator: n·Σxy − Σx·Σy = 8(357,550) − 6150(434) = 2,860,400 − 2,669,100 = 191,300
    
    Denominator:
    - n·Σx² − (Σx)² = 42,820,000 − 37,822,500 = 4,997,500
    - n·Σy² − (Σy)² = 196,128 − 188,356 = 7,772
    - √(4,997,500 × 7,772) ≈ 197,080.1
    
    r = 191,300 / 197,080.1 ≈ 0.9707
    
    Final answer: 0.97

    492 tokens in, 1,515 out (1,302 of them reasoning) · 4.8 s · $0.0013 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Novita ·

  • Claude Opus 5.5

    Passed: Final answer 0.97: right.

    **Means**
    - Mean ad_spend: 6150 / 8 = 768.75
    - Mean signups: 434 / 8 = 54.25
    
    **Sums of deviations**
    - Σ(dx·dy) = 23,912.5
    - Σ(dx²) = 624,687.5
    - Σ(dy²) = 971.5
    
    **Pearson r**
    
    r = 23,912.5 / √(624,687.5 × 971.5)
    = 23,912.5 / √606,883,906.25
    = 23,912.5 / 24,635.01
    ≈ 0.9707
    
    This is a strong positive correlation.
    
    Final answer: 0.97

    607 tokens in, 1,155 out (946 of them reasoning) · 11.5 s · $0.0261 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  • Claude Sonnet 5

    Failed: Final answer 0.99; expected 0.97.

    **Working:**
    
    n=8, mean(x)=768.75, mean(y)=54.25
    
    Computed deviations and cross-products:
    - Σ(dx·dy) = 24428.125
    - Σ(dx²) = 624687.5
    - Σ(dy²) = 971.5
    
    r = Σ(dx·dy) / √(Σdx² · Σdy²) = 24428.125 / √(624687.5 × 971.5) = 24428.125 / 24635.01 ≈ 0.9916
    
    **Final answer: 0.99**

    605 tokens in, 1,524 out (1,348 of them reasoning) · 13.8 s · $0.0165 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

How we ranked them

Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand.

Questions

What is the best AI model right now?

In our own test runs on September 2026, DeepSeek V4.1 Flash, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 49 of 50, the most; DeepSeek V4.1 Flash is first on more of the hard ones (20 of 20). The top 6 models passed between 47 and 49, so one prompt can move a model several places: the best one for you depends on your job, which the table by job shows.

What is the smartest AI?

On our hard prompts alone, DeepSeek V4.1 Flash, GPT-6 Astra, GPT-6 Luna, and Grok 4.7 passed 20 of 20 each. "Smartest" here means the hardest prompts we set, from multi-step math to code with edge cases; it's one test, not a verdict on intelligence.

Which AI is best: ChatGPT, Claude or Gemini?

We tested the models, not the apps. Of OpenAI's, GPT-6 Astra passed 48 of 50; of Anthropic's, Claude Opus 5.5 passed 49 of 50; of Google's, Gemini 3.1 Pro passed 49 of 50. ChatGPT, Claude and Gemini add their own tools, limits and settings on top.

How often does this ranking change?

Each time we run the same prompts again, and a model that joins llmwise is ranked once it has been run. The date of the runs is at the top of this page. Runs count for 45 days; after that, this page says they're out of date until we run the prompts again.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.