Skip to content

Best AI · Business

The best AI for business, from our test runs

On the business jobs we test (writing, replying to customers, working out numbers from a table, and summaries), DeepSeek V4 Pro, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 19 of 20, the most; DeepSeek V4 Pro is first on more of the hard ones (8 of 8). Every model ranked, the best at each price, and what the test doesn't cover.

Test runs and prices checked . Updated .

Short answer

DeepSeek V4 Pro, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 19 of 20, the most; DeepSeek V4 Pro is first on more of the hard ones (8 of 8); the best everyday model was GPT-6 Luna, 18 of 20. This is our own test, run on October 9, 2026: every prompt, reply and score is published.

Ranked by our test runs

This is our own test: the same prompts sent to every model through llmwise, each reply checked the same way, with every prompt, reply and score published. The business ranking adds up four jobs from our test set: writing, customer support, data analysis, and summarization. It doesn't test your own tools, such as a CRM or accounts software.

Models ranked on writing, customer support, data analysis, and summarization
#ModelPassedHard onesFree trialOn ProCost per reply
1DeepSeek V4 ProDeepSeek19 of 208 of 8Yes250 a month$0.0016
2Gemini 3.1 Pro (preview)Google19 of 207 of 8Yes125 a month$0.0110
3Claude Opus 5.5Anthropic19 of 207 of 81 message62 a month$0.0130
4GPT-6 LunaOpenAI18 of 208 of 8Yes60 a day$0.0001
5GPT-6 AstraOpenAI18 of 208 of 8No31 a month$0.0123
6DeepSeek V4.1 FlashDeepSeek18 of 207 of 8Yes60 a day$0.0005
7GLM 5.3Z.ai18 of 207 of 8Yes250 a month$0.0009
8Mistral Large 4Mistral18 of 206 of 8Yes250 a month$0.0030
9Grok 4.7xAI17 of 207 of 8Yes250 a month$0.0065
10Claude Sonnet 5.5Anthropic17 of 207 of 8Yes125 a month$0.0040
11GPT-6 SolOpenAI16 of 208 of 8Yes125 a month$0.0027
12GPT-6.1 SolOpenAI16 of 207 of 8Yes125 a month$0.0012
13Claude Sonnet 5Anthropic16 of 207 of 8Yes125 a month$0.0042
14Gemini 3.8 FlashGoogle16 of 206 of 8Yes250 a month$0.0016
15Kimi K3Moonshot16 of 206 of 8Yes125 a month$0.0047
16Claude Haiku 5.5Anthropic15 of 207 of 8Yes60 a day$0.0003
17Claude Haiku 4.5Anthropic15 of 205 of 8Yes250 a month$0.0015
18Claude Fable 5.1Anthropic15 of 205 of 8No31 a month$0.0209
19GLM 5.3 FlashZ.ai11 of 204 of 8Yes60 a day$0.0003
20 prompts per model (writing, customer support, data analysis, and summarization), run on October 9, 2026. Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand. Cost per reply is what OpenRouter charged us on average; in llmwise you pay per message.

The prompts, every reply and how each was scored: our test runs.

The best pick at each price

The same results, by what a message counts as on Pro: the best model at each size of message, then the rest of that size.

  • 250 a month on Pro

    DeepSeek V4 Pro, 19 of 20 passed; then GLM 5.3 18 of 20, Mistral Large 4 18 of 20, Grok 4.7 17 of 20, Gemini 3.8 Flash 16 of 20, and Claude Haiku 4.5 15 of 20.

  • 125 a month on Pro

    Gemini 3.1 Pro, 19 of 20 passed; then Claude Sonnet 5.5 17 of 20, GPT-6 Sol 16 of 20, GPT-6.1 Sol 16 of 20, Claude Sonnet 5 16 of 20, and Kimi K3 16 of 20.

  • 62 a month on Pro

    Claude Opus 5.5, 19 of 20 passed.

  • 60 a day on Pro (everyday models)

    GPT-6 Luna, 18 of 20 passed; then DeepSeek V4.1 Flash 18 of 20, Claude Haiku 5.5 15 of 20, and GLM 5.3 Flash 11 of 20.

  • 31 a month on Pro

    GPT-6 Astra, 18 of 20 passed; then Claude Fable 5.1 15 of 20.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Where the top models differ

The top 5, job by job.

Prompts passed on each job, by model
JobDeepSeek V4 ProGemini 3.1 ProClaude Opus 5.5GPT-6 LunaGPT-6 Astra
Writing5 of 54 of 54 of 55 of 55 of 5
Customer support4 of 55 of 55 of 54 of 53 of 5
Data analysis5 of 55 of 55 of 55 of 55 of 5
Summarization5 of 55 of 55 of 54 of 55 of 5

One prompt, side by side

An everyday customer support prompt, “A frustrated customer”, where the models didn't all pass.

The prompt: A frustrated customer (everyday)
Northwind Outfitters support policy
- Returns: within 30 days of delivery, unworn items with tags get a full refund to the original payment method.
- After 30 days and up to 60 days: exchange or a store gift card only, no refund.
- Refunds are paid within 5 business days of the return reaching our warehouse.
- Standard delivery takes 3 to 5 business days. If an order is more than 7 business days late, the customer gets free express shipping on their next order.
- Staff can't give discount codes worth more than 15%.

Customer message: "This is the THIRD time you've sent the wrong size. I ordered a medium and got a large AGAIN. I'm so fed up with this."

Write the reply to the customer, following the policy. At most 150 words, signed "Sam, Northwind support".
  • Claude Opus 5.5

    Passed: Graded 4.0 of 5 on average (lowest 3).

    Hi there,
    
    I'm really sorry. Receiving the wrong size once is frustrating, and three times is unacceptable. You ordered a medium, and that's what you should have received.
    
    Here's how we can fix this:
    
    - **Exchange for the correct size:** Send the large back and we'll get a medium out to you.
    - **Full refund:** If you'd rather not wait, returns within 30 days of delivery get a full refund to your original payment method. Refunds are paid within 5 business days of the item reaching our warehouse.
    
    Please reply with your order number and which option you prefer. I'll also flag this repeated error with our warehouse team so it doesn't happen again.
    
    As an apology, I'd like to offer you a 15% discount code for your next order.
    
    Thank you for your patience. I know it's been tested.
    
    Sam, Northwind support

    720 tokens in, 376 out (109 of them reasoning) · 6.4 s · $0.0111 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  • DeepSeek V4.1 Flash

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Hi — I'm truly sorry. Three wrong sizes in a row is not the standard we hold ourselves to, and I understand your frustration.
    
    Here's what I'll do right now:
    
    - I've flagged your order for a size check before it ships again. Please confirm you need a **medium**, and I'll send the correct item at no cost.
    - A prepaid return label is on its way for the large, so there's nothing for you to pay.
    - As an apology, here's a 15% code for your next order: **SORRY15**.
    
    On refunds, just so you have it: unworn items with tags returned within 30 days of delivery are refunded in full to your original payment method, usually within 5 business days of reaching our warehouse.
    
    Reply with your order number and I'll handle the rest personally.
    
    Sam, Northwind support

    561 tokens in, 336 out (155 of them reasoning) · 1.3 s · $0.0005 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Together ·

  • GPT-6 Sol

    Failed: Graded 3.0 of 5 on average (lowest 2).

    I’m sorry we sent the wrong size again—especially for the third time. You ordered a medium, and you should not have to keep chasing us to get it right.
    
    Please send me your order number so I can check what happened and help arrange an exchange for the medium. If you’d rather return it, an unworn item with tags is eligible for a full refund to your original payment method within 30 days of delivery. After 30 days and up to 60 days, we can offer an exchange or store gift card.
    
    Sam, Northwind support

    516 tokens in, 324 out (205 of them reasoning) · 5.9 s · $0.0043 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

How we ranked them

Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand.

Questions

What is the best AI for business?

In our own test runs on October 2026, on the business prompts (writing, customer support, data analysis, and summarization): DeepSeek V4 Pro, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 19 of 20, the most; DeepSeek V4 Pro is first on more of the hard ones (8 of 8); of the everyday models, GPT-6 Luna did best (18 of 20), at up to 60 messages a day on Pro.

Where do AI models go wrong on business work?

Replying to customers split the models most: on the prompt with a frustrated customer, 8 of 19 models passed. Most failed by promising something the policy didn't allow or by missing what the customer asked.

Did you test Microsoft Copilot or Gemini for Workspace?

No. We ran the models llmwise offers, on our own prompts, through llmwise. Business apps built on AI add their own tools, data access and limits, which we didn't test.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.