Best AI · Business
The best AI for business, from our test runs
On the business jobs we test (writing, replying to customers, working out numbers from a table, and summaries), DeepSeek V4 Pro, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 19 of 20, the most; DeepSeek V4 Pro is first on more of the hard ones (8 of 8). Every model ranked, the best at each price, and what the test doesn't cover.
Test runs and prices checked . Updated .
Short answer
DeepSeek V4 Pro, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 19 of 20, the most; DeepSeek V4 Pro is first on more of the hard ones (8 of 8); the best everyday model was GPT-6 Luna, 18 of 20. This is our own test, run on October 9, 2026: every prompt, reply and score is published.
Ranked by our test runs
This is our own test: the same prompts sent to every model through llmwise, each reply checked the same way, with every prompt, reply and score published. The business ranking adds up four jobs from our test set: writing, customer support, data analysis, and summarization. It doesn't test your own tools, such as a CRM or accounts software.
| # | Model | Passed | Hard ones | Free trial | On Pro | Cost per reply |
|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 ProDeepSeek | 19 of 20 | 8 of 8 | Yes | 250 a month | $0.0016 |
| 2 | Gemini 3.1 Pro (preview)Google | 19 of 20 | 7 of 8 | Yes | 125 a month | $0.0110 |
| 3 | Claude Opus 5.5Anthropic | 19 of 20 | 7 of 8 | 1 message | 62 a month | $0.0130 |
| 4 | GPT-6 LunaOpenAI | 18 of 20 | 8 of 8 | Yes | 60 a day | $0.0001 |
| 5 | GPT-6 AstraOpenAI | 18 of 20 | 8 of 8 | No | 31 a month | $0.0123 |
| 6 | DeepSeek V4.1 FlashDeepSeek | 18 of 20 | 7 of 8 | Yes | 60 a day | $0.0005 |
| 7 | GLM 5.3Z.ai | 18 of 20 | 7 of 8 | Yes | 250 a month | $0.0009 |
| 8 | Mistral Large 4Mistral | 18 of 20 | 6 of 8 | Yes | 250 a month | $0.0030 |
| 9 | Grok 4.7xAI | 17 of 20 | 7 of 8 | Yes | 250 a month | $0.0065 |
| 10 | Claude Sonnet 5.5Anthropic | 17 of 20 | 7 of 8 | Yes | 125 a month | $0.0040 |
| 11 | GPT-6 SolOpenAI | 16 of 20 | 8 of 8 | Yes | 125 a month | $0.0027 |
| 12 | GPT-6.1 SolOpenAI | 16 of 20 | 7 of 8 | Yes | 125 a month | $0.0012 |
| 13 | Claude Sonnet 5Anthropic | 16 of 20 | 7 of 8 | Yes | 125 a month | $0.0042 |
| 14 | Gemini 3.8 FlashGoogle | 16 of 20 | 6 of 8 | Yes | 250 a month | $0.0016 |
| 15 | Kimi K3Moonshot | 16 of 20 | 6 of 8 | Yes | 125 a month | $0.0047 |
| 16 | Claude Haiku 5.5Anthropic | 15 of 20 | 7 of 8 | Yes | 60 a day | $0.0003 |
| 17 | Claude Haiku 4.5Anthropic | 15 of 20 | 5 of 8 | Yes | 250 a month | $0.0015 |
| 18 | Claude Fable 5.1Anthropic | 15 of 20 | 5 of 8 | No | 31 a month | $0.0209 |
| 19 | GLM 5.3 FlashZ.ai | 11 of 20 | 4 of 8 | Yes | 60 a day | $0.0003 |
The prompts, every reply and how each was scored: our test runs.
The best pick at each price
The same results, by what a message counts as on Pro: the best model at each size of message, then the rest of that size.
250 a month on Pro
DeepSeek V4 Pro, 19 of 20 passed; then GLM 5.3 18 of 20, Mistral Large 4 18 of 20, Grok 4.7 17 of 20, Gemini 3.8 Flash 16 of 20, and Claude Haiku 4.5 15 of 20.
125 a month on Pro
Gemini 3.1 Pro, 19 of 20 passed; then Claude Sonnet 5.5 17 of 20, GPT-6 Sol 16 of 20, GPT-6.1 Sol 16 of 20, Claude Sonnet 5 16 of 20, and Kimi K3 16 of 20.
62 a month on Pro
Claude Opus 5.5, 19 of 20 passed.
60 a day on Pro (everyday models)
GPT-6 Luna, 18 of 20 passed; then DeepSeek V4.1 Flash 18 of 20, Claude Haiku 5.5 15 of 20, and GLM 5.3 Flash 11 of 20.
31 a month on Pro
GPT-6 Astra, 18 of 20 passed; then Claude Fable 5.1 15 of 20.
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
Where the top models differ
The top 5, job by job.
| Job | DeepSeek V4 Pro | Gemini 3.1 Pro | Claude Opus 5.5 | GPT-6 Luna | GPT-6 Astra |
|---|---|---|---|---|---|
| Writing | 5 of 5 | 4 of 5 | 4 of 5 | 5 of 5 | 5 of 5 |
| Customer support | 4 of 5 | 5 of 5 | 5 of 5 | 4 of 5 | 3 of 5 |
| Data analysis | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Summarization | 5 of 5 | 5 of 5 | 5 of 5 | 4 of 5 | 5 of 5 |
One prompt, side by side
An everyday customer support prompt, “A frustrated customer”, where the models didn't all pass.
The prompt: A frustrated customer (everyday)
Northwind Outfitters support policy - Returns: within 30 days of delivery, unworn items with tags get a full refund to the original payment method. - After 30 days and up to 60 days: exchange or a store gift card only, no refund. - Refunds are paid within 5 business days of the return reaching our warehouse. - Standard delivery takes 3 to 5 business days. If an order is more than 7 business days late, the customer gets free express shipping on their next order. - Staff can't give discount codes worth more than 15%. Customer message: "This is the THIRD time you've sent the wrong size. I ordered a medium and got a large AGAIN. I'm so fed up with this." Write the reply to the customer, following the policy. At most 150 words, signed "Sam, Northwind support".
Claude Opus 5.5
Passed: Graded 4.0 of 5 on average (lowest 3).
Hi there, I'm really sorry. Receiving the wrong size once is frustrating, and three times is unacceptable. You ordered a medium, and that's what you should have received. Here's how we can fix this: - **Exchange for the correct size:** Send the large back and we'll get a medium out to you. - **Full refund:** If you'd rather not wait, returns within 30 days of delivery get a full refund to your original payment method. Refunds are paid within 5 business days of the item reaching our warehouse. Please reply with your order number and which option you prefer. I'll also flag this repeated error with our warehouse team so it doesn't happen again. As an apology, I'd like to offer you a 15% discount code for your next order. Thank you for your patience. I know it's been tested. Sam, Northwind support
720 tokens in, 376 out (109 of them reasoning) · 6.4 s · $0.0111 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·
DeepSeek V4.1 Flash
Passed: Graded 4.7 of 5 on average (lowest 4).
Hi — I'm truly sorry. Three wrong sizes in a row is not the standard we hold ourselves to, and I understand your frustration. Here's what I'll do right now: - I've flagged your order for a size check before it ships again. Please confirm you need a **medium**, and I'll send the correct item at no cost. - A prepaid return label is on its way for the large, so there's nothing for you to pay. - As an apology, here's a 15% code for your next order: **SORRY15**. On refunds, just so you have it: unworn items with tags returned within 30 days of delivery are refunded in full to your original payment method, usually within 5 business days of reaching our warehouse. Reply with your order number and I'll handle the rest personally. Sam, Northwind support
561 tokens in, 336 out (155 of them reasoning) · 1.3 s · $0.0005 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Together ·
GPT-6 Sol
Failed: Graded 3.0 of 5 on average (lowest 2).
I’m sorry we sent the wrong size again—especially for the third time. You ordered a medium, and you should not have to keep chasing us to get it right. Please send me your order number so I can check what happened and help arrange an exchange for the medium. If you’d rather return it, an unworn item with tags is eligible for a full refund to your original payment method within 30 days of delivery. After 30 days and up to 60 days, we can offer an exchange or store gift card. Sam, Northwind support
516 tokens in, 324 out (205 of them reasoning) · 5.9 s · $0.0043 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
How we ranked them
Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand.
- Our test runs: 50 prompts on every model
- The AI model leaderboard, job by job
- Best AI assistants, tested (October 2026)
- The best AI for research, from our test runs
- The best AI for students, from our test runs
- The best free AI, from our test runs
- The best AI model right now: all 19, ranked by our tests
- The best AI for data analysis, from our test runs
- AI chat with every top model in one place
Questions
What is the best AI for business?
In our own test runs on October 2026, on the business prompts (writing, customer support, data analysis, and summarization): DeepSeek V4 Pro, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 19 of 20, the most; DeepSeek V4 Pro is first on more of the hard ones (8 of 8); of the everyday models, GPT-6 Luna did best (18 of 20), at up to 60 messages a day on Pro.
Where do AI models go wrong on business work?
Replying to customers split the models most: on the prompt with a frustrated customer, 8 of 19 models passed. Most failed by promising something the policy didn't allow or by missing what the customer asked.
Did you test Microsoft Copilot or Gemini for Workspace?
No. We ran the models llmwise offers, on our own prompts, through llmwise. Business apps built on AI add their own tools, data access and limits, which we didn't test.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.