Skip to content

Best AI · Data analysis

The best AI for data analysis, from our test runs

We gave every model the same small tables of orders and ad spend and asked questions with one right answer: a top region, an average, a change, a filtered median and a correlation. GPT-6 Luna, Claude Haiku 5.5, DeepSeek V4.1 Flash, GLM 5.3, DeepSeek V4 Pro, Mistral Large 4, Grok 4.7, GPT-6.1 Sol, GPT-6 Sol, Claude Sonnet 5.5, Kimi K3, Gemini 3.1 Pro, Claude Opus 5.5, GPT-6 Astra, and Claude Fable 5.1 each passed 5 of 5, the most; GPT-6 Luna is first on the tie-break (the smaller message, then the lower cost).

Test runs and prices checked . Updated .

Short answer

GPT-6 Luna, Claude Haiku 5.5, DeepSeek V4.1 Flash, GLM 5.3, DeepSeek V4 Pro, Mistral Large 4, Grok 4.7, GPT-6.1 Sol, GPT-6 Sol, Claude Sonnet 5.5, Kimi K3, Gemini 3.1 Pro, Claude Opus 5.5, GPT-6 Astra, and Claude Fable 5.1 each passed 5 of 5, the most; GPT-6 Luna is first on the tie-break (the smaller message, then the lower cost). This is our own test, run on October 9, 2026: every prompt, reply and score is published.

Ranked by our test runs

This is our own test: the same prompts sent to every model through llmwise, each reply checked the same way, with every prompt, reply and score published. The data prompts give each model a small table as CSV and check its final number; they don't test charts, spreadsheet formulas or large files.

Models ranked on data analysis
#ModelPassedHard onesFree trialOn ProCost per reply
1GPT-6 LunaOpenAI5 of 52 of 2Yes60 a day$0.0002
2Claude Haiku 5.5Anthropic5 of 52 of 2Yes60 a day$0.0004
3DeepSeek V4.1 FlashDeepSeek5 of 52 of 2Yes60 a day$0.0008
4GLM 5.3Z.ai5 of 52 of 2Yes250 a month$0.0019
5DeepSeek V4 ProDeepSeek5 of 52 of 2Yes250 a month$0.0029
6Mistral Large 4Mistral5 of 52 of 2Yes250 a month$0.0037
7Grok 4.7xAI5 of 52 of 2Yes250 a month$0.0111
8GPT-6.1 SolOpenAI5 of 52 of 2Yes125 a month$0.0016
9GPT-6 SolOpenAI5 of 52 of 2Yes125 a month$0.0034
10Claude Sonnet 5.5Anthropic5 of 52 of 2Yes125 a month$0.0047
11Kimi K3Moonshot5 of 52 of 2Yes125 a month$0.0068
12Gemini 3.1 Pro (preview)Google5 of 52 of 2Yes125 a month$0.0155
13Claude Opus 5.5Anthropic5 of 52 of 21 message62 a month$0.0131
14GPT-6 AstraOpenAI5 of 52 of 2No31 a month$0.0154
15Claude Fable 5.1Anthropic5 of 52 of 2No31 a month$0.0283
16GLM 5.3 FlashZ.ai4 of 51 of 2Yes60 a day$0.0004
17Gemini 3.8 FlashGoogle4 of 51 of 2Yes250 a month$0.0028
18Claude Sonnet 5Anthropic4 of 51 of 2Yes125 a month$0.0069
19Claude Haiku 4.5Anthropic3 of 51 of 2Yes250 a month$0.0025
5 prompts per model (data analysis), run on October 9, 2026. Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand. Cost per reply is what OpenRouter charged us on average; in llmwise you pay per message.

The prompts, every reply and how each was scored: our test runs.

The best pick at each price

The same results, by what a message counts as on Pro: the best model at each size of message, then the rest of that size.

  • 60 a day on Pro (everyday models)

    GPT-6 Luna, 5 of 5 passed; then Claude Haiku 5.5 5 of 5, DeepSeek V4.1 Flash 5 of 5, and GLM 5.3 Flash 4 of 5.

  • 250 a month on Pro

    GLM 5.3, 5 of 5 passed; then DeepSeek V4 Pro 5 of 5, Mistral Large 4 5 of 5, Grok 4.7 5 of 5, Gemini 3.8 Flash 4 of 5, and Claude Haiku 4.5 3 of 5.

  • 125 a month on Pro

    GPT-6.1 Sol, 5 of 5 passed; then GPT-6 Sol 5 of 5, Claude Sonnet 5.5 5 of 5, Kimi K3 5 of 5, Gemini 3.1 Pro 5 of 5, and Claude Sonnet 5 4 of 5.

  • 62 a month on Pro

    Claude Opus 5.5, 5 of 5 passed.

  • 31 a month on Pro

    GPT-6 Astra, 5 of 5 passed; then Claude Fable 5.1 5 of 5.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

One prompt, side by side

An everyday data analysis prompt, “Average order value in August”, where the models didn't all pass.

The prompt: Average order value in August (everyday)
Here are some orders as CSV. Revenue is quantity times price. What was the average revenue per order in August 2026, to 2 decimal places?

order_id,date,region,product,quantity,price
1001,2026-07-03,West,Lamp,2,45.00
1002,2026-07-05,East,Chair,1,120.00
1003,2026-07-09,West,Desk,1,310.00
1004,2026-07-14,North,Lamp,4,45.00
1005,2026-07-21,East,Lamp,3,45.00
1006,2026-07-28,North,Chair,2,120.00
1007,2026-08-02,West,Chair,2,120.00
1008,2026-08-06,East,Desk,1,310.00
1009,2026-08-11,North,Desk,2,310.00
1010,2026-08-15,West,Lamp,5,45.00
1011,2026-08-19,East,Chair,3,120.00
1012,2026-08-27,North,Lamp,1,45.00

Show your working briefly, then end with a line that says "Final answer: " followed by the answer alone.
  • GPT-6 Luna

    Passed: Final answer 300.00: right.

    August revenue: $240 + $310 + $620 + $225 + $360 + $45 = $1,800.  
    There were 6 orders, so the average is $1,800 ÷ 6 = $300.00.
    
    Final answer: $300.00

    647 tokens in, 131 out (67 of them reasoning) · 2.6 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  • Claude Sonnet 5.5

    Passed: Final answer 300.00: right.

    August 2026 orders (1007–1012):
    
    - 1007: 2 × 120 = 240
    - 1008: 1 × 310 = 310
    - 1009: 2 × 310 = 620
    - 1010: 5 × 45 = 225
    - 1011: 3 × 120 = 360
    - 1012: 1 × 45 = 45
    
    Total = 240 + 310 + 620 + 225 + 360 + 45 = 1800
    
    Average = 1800 / 6 = 300.00
    
    Final answer: 300.00

    825 tokens in, 170 out · 1.7 s · $0.0037 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·

  • Claude Haiku 4.5

    Failed: Final answer 275.00; expected 300.

    I'll find the average revenue per order for August 2026.
    
    **August 2026 orders:**
    
    | order_id | date | quantity | price | revenue |
    |----------|------|----------|-------|---------|
    | 1007 | 2026-08-02 | 2 | 45.00 | 90.00 |
    | 1008 | 2026-08-06 | 1 | 310.00 | 310.00 |
    | 1009 | 2026-08-11 | 2 | 310.00 | 620.00 |
    | 1010 | 2026-08-15 | 5 | 45.00 | 225.00 |
    | 1011 | 2026-08-19 | 3 | 120.00 | 360.00 |
    | 1012 | 2026-08-27 | 1 | 45.00 | 45.00 |
    
    **Calculation:**
    - Total revenue: 90 + 310 + 620 + 225 + 360 + 45 = 1,650.00
    - Number of orders: 6
    - Average revenue per order: 1,650.00 ÷ 6 = 275.00
    
    Final answer: 275.00

    683 tokens in, 326 out · 2.1 s · $0.0023 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

How we ranked them

Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand.

Questions

What is the best AI for data analysis?

In our own test runs on October 2026: GPT-6 Luna, Claude Haiku 5.5, DeepSeek V4.1 Flash, GLM 5.3, DeepSeek V4 Pro, Mistral Large 4, Grok 4.7, GPT-6.1 Sol, GPT-6 Sol, Claude Sonnet 5.5, Kimi K3, Gemini 3.1 Pro, Claude Opus 5.5, GPT-6 Astra, and Claude Fable 5.1 each passed 5 of 5, the most; GPT-6 Luna is first on the tie-break (the smaller message, then the lower cost); of the everyday models, GPT-6 Luna did best (5 of 5). Each answer is checked against the right number, so a pass means the arithmetic was right.

Is this the best AI for Excel?

Partly. These prompts test working out answers from a table (totals, averages, a filtered median, a correlation), which is the heart of spreadsheet analysis. They don't test writing Excel formulas or macros, so we don't rank models for that.

Can the AI run code on my data?

On paid plans the model can run code on a file you attach and show the code it ran. These runs don't use it: every model got the table in the message and had to work it out.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.