Skip to content

Best AI · 2026

The best AI chatbots in 2026, and when to use each

19 models from 8 AI companies, Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, priced per message from Claude Haiku 5.5, up to 60 messages a day on Pro, to GPT-6 Astra, up to 31 messages a month. What each is for, what a message costs, and which to reach for when.

Based on 950 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Our picks, job by job

From our test runs: the same 10 jobs' prompts on every model, each pick by fixed rules from the results. Each job's page has every reply.

Our picks for each job
JobHard problemsBest valueEveryday
CodingShared by 17 models (5 of 5)GLM 5.3 (5 of 5)GLM 5.3 Flash (5 of 5)
WritingShared by 3 models (5 of 5)DeepSeek V4 Pro (5 of 5)GPT-6 Luna (5 of 5)
MathShared by 18 models (5 of 5)GLM 5.3 (5 of 5)Claude Haiku 5.5 (5 of 5)
SummarizationShared by 10 models (5 of 5)Gemini 3.8 Flash (5 of 5)DeepSeek V4.1 Flash (5 of 5)
Data analysisShared by 15 models (5 of 5)GLM 5.3 (5 of 5)GPT-6 Luna (5 of 5)
Customer supportShared by 4 models (5 of 5)GLM 5.3 (5 of 5)GPT-6 Luna (4 of 5)
TranslationShared by 18 models (5 of 5)GLM 5.3 (5 of 5)GPT-6 Luna (5 of 5)
SQLShared by 19 models (5 of 5)GLM 5.3 (5 of 5)GPT-6 Luna (5 of 5)
RAG and answering from documentsShared by 19 models (5 of 5)GLM 5.3 (5 of 5)GPT-6 Luna (5 of 5)
Agents and tool useShared by 18 models (5 of 5)GLM 5.3 (5 of 5)GPT-6 Luna (5 of 5)
Each job has five prompts. The picks follow the same published rules on every page, from the results alone.

Every model ranked in one list, and on the hard prompts alone: the best AI model right now.

Does the best AI cost the most?

Not on our prompts. 3 models shared the top result in our runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026, 49 of 50: DeepSeek V4 Pro, Gemini 3.1 Pro (preview), Claude Opus 5.5. A reply cost $0.0037 on DeepSeek V4 Pro and $0.0107 on Claude Opus 5.5, about 3 times as much.

Fifty prompts don't settle which model is best for your work; they do show that price and quality aren't the same axis. The makers publish what each model costs and how much it reads at once: Anthropic's models overview and OpenAI's GPT-6 Astra page.

The models, family by family

  • Claude by Anthropic

    Claude is the family of AI models made by Anthropic.

    • Claude Fable 5.1 (31/mo on Pro) — Anthropic's most capable model for the hardest problems.
    • Claude Opus 5.5 (62/mo on Pro) — Deep reasoning for complex work.
    • Claude Sonnet 5.5 (125/mo on Pro) — Newest Sonnet, for coding, documents and everyday work.
    • Claude Sonnet 5 (125/mo on Pro) — Fast, smart all-rounder for writing and reasoning.
    • Claude Haiku 4.5 (250/mo on Pro) — Quick answers at low cost.
    • Claude Haiku 5.5 (60/day on Pro) — Newest Haiku, fast and low-cost, with thinking.
  • GPT by OpenAI

    GPT is the family of AI models made by OpenAI, the company behind ChatGPT.

    • GPT-6 Astra (31/mo on Pro) — OpenAI's flagship model.
    • GPT-6.1 Sol (125/mo on Pro) — Newest Sol, near-Astra for coding and agentic work.
    • GPT-6 Sol (125/mo on Pro) — Strong reasoning for coding and agentic work.
    • GPT-6 Luna (60/day on Pro) — Small, fast, and inexpensive.
  • Gemini by Google

    Gemini is Google's family of AI models.

    • Gemini 3.1 Pro (preview) (125/mo on Pro) — Google's most capable model (preview).
    • Gemini 3.8 Flash (250/mo on Pro) — Fast multimodal model.
  • DeepSeek by DeepSeek

    DeepSeek's models are made by the AI company of the same name.

    • DeepSeek V4 Pro (250/mo on Pro) — DeepSeek's most capable model.
    • DeepSeek V4.1 Flash (60/day on Pro) — Fast, low-cost DeepSeek model with vision.
  • Grok by xAI

    Grok is the family of AI models made by xAI.

    • Grok 4.7 (250/mo on Pro) — xAI's flagship for coding and agentic work.
  • Kimi by Moonshot

    Kimi is the family of AI models made by Moonshot.

    • Kimi K3 (125/mo on Pro) — Moonshot's long-context reasoning model.
  • GLM by Z.ai

    GLM is the family of AI models made by Z.ai.

    • GLM 5.3 (250/mo on Pro) — Z.ai's capable open-weight model.
    • GLM 5.3 Flash (60/day on Pro) — Fast, low-cost GLM model with vision.
  • Mistral by Mistral

    Mistral's models are made by Mistral AI, a French AI company.

    • Mistral Large 4 (250/mo on Pro) — Mistral's flagship for reasoning, coding and agentic work.

Which model for what

Suggestions by price band, not rankings.

  • Everyday models

    For quick questions, rewrites, lookups and anything you ask often.

    • Claude Haiku 5.560/day on Pro
    • GPT-6 Luna60/day on Pro
    • DeepSeek V4.1 Flash60/day on Pro
    • GLM 5.3 Flash60/day on Pro
  • Models with 125 to 250 messages a month on Pro

    For most real work: writing, explaining, analysis and code, where a better answer is worth one of your monthly messages.

    • Claude Sonnet 5.5125/mo on Pro
    • Claude Sonnet 5125/mo on Pro
    • Claude Haiku 4.5250/mo on Pro
    • GPT-6.1 Sol125/mo on Pro
    • GPT-6 Sol125/mo on Pro
    • Gemini 3.1 Pro (preview)125/mo on Pro
    • Gemini 3.8 Flash250/mo on Pro
    • DeepSeek V4 Pro250/mo on Pro
    • Grok 4.7250/mo on Pro
    • Kimi K3125/mo on Pro
    • GLM 5.3250/mo on Pro
    • Mistral Large 4250/mo on Pro
  • Models with 31 to 62 messages a month on Pro

    For the hardest problems, where getting it right matters more than the price.

    • Claude Fable 5.131/mo on Pro
    • Claude Opus 5.562/mo on Pro
    • GPT-6 Astra31/mo on Pro

The best AI for each job

Each job has its own page: what matters, which models fit, and tips.

  • Agents and tool use

    Look for: reliable tool calls, following a plan, knowing when to ask, and room for tool output.

  • Coding

    Look for: room for your code, reasoning on hard problems, running the code, and cost per try.

  • Customer support

    Look for: tone, accuracy about your product, cost at volume, and languages.

  • Data analysis

    Look for: calculations in code, reading pdfs and charts, and explaining the result.

  • Math

    Look for: step-by-step reasoning, checking the arithmetic, and showing the work.

  • RAG and answering from documents

    Look for: room for the documents, sticking to the sources, and citations.

  • SQL

    Look for: knowing your schema, the right dialect, and correctness.

  • Summarization

    Look for: how much it can read, staying faithful, and the right length.

  • Translation

    Look for: your language pair, tone and register, and keeping the formatting.

  • Writing

    Look for: voice and tone, following the brief, and editing, not just drafting.

Every model at a glance

Every model in llmwise
ModelOn ProOn FreeContext windowImagesPDFsReasoningAPI price per 1M, in / out
Claude Fable 5.1Anthropic31/mo on ProNo1M tokensYesWhole fileYes$10.00 / $50.00
Claude Opus 5.5Anthropic62/mo on Pro1 message1M tokensYesWhole fileYes$4.00 / $20.00
Claude Sonnet 5.5Anthropic125/mo on ProYes1M tokensYesWhole fileYes$2.00 / $10.00
Claude Sonnet 5Anthropic125/mo on ProYes1M tokensYesWhole fileYes$2.00 / $10.00
Claude Haiku 5.5Anthropic60/day on ProYes1M tokensYesWhole fileYes$0.10 / $0.50
Claude Haiku 4.5Anthropic250/mo on ProYes200K tokensYesWhole fileNo$1.00 / $5.00
GPT-6 AstraOpenAI31/mo on ProNo1.05M tokensYesWhole fileYes$10.00 / $50.00
GPT-6.1 SolOpenAI125/mo on ProYes1.05M tokensYesWhole fileYes$2.00 / $10.00
GPT-6 SolOpenAI125/mo on ProYes1.05M tokensYesWhole fileYes$2.00 / $10.00
GPT-6 LunaOpenAI60/day on ProYes1.05M tokensYesWhole fileYes$0.10 / $0.50
Gemini 3.1 Pro (preview)Google125/mo on ProYes1.05M tokensYesWhole fileYes$2.00 / $12.00
Gemini 3.8 FlashGoogle250/mo on ProYes1.05M tokensYesWhole fileYes$0.75 / $3.75
DeepSeek V4.1 FlashDeepSeek60/day on ProYes1.05M tokensYesText onlyYes$0.30 / $1.20
DeepSeek V4 ProDeepSeek250/mo on ProYes1.05M tokensNoText onlyYes$0.40 / $4.00
Grok 4.7xAI250/mo on ProYes500K tokensYesWhole fileYes$2.00 / $6.00
Kimi K3Moonshot125/mo on ProYes1.05M tokensYesText onlyYes$3.00 / $15.00
GLM 5.3Z.ai250/mo on ProYes1.05M tokensNoText onlyYes$1.40 / $4.40
GLM 5.3 FlashZ.ai60/day on ProYes1.05M tokensYesText onlyYes$0.15 / $0.50
Mistral Large 4Mistral250/mo on ProYes1.05M tokensYesText onlyYes$0.68 / $2.09
Each badge is how many messages Pro gets on the model: a month’s, or a day’s on an everyday model. Free is a one-time trial of 5 messages on the models marked. “Text only” models get the text of a PDF, not the file. API prices are the per-token prices in our model catalog as of October 2026 (Anthropic: Anthropic's list price; OpenAI: OpenAI's list price; Google: Google's list price; DeepSeek: the price of the OpenRouter endpoints llmwise uses, not DeepSeek's own API; xAI: xAI's price, served through OpenRouter; Moonshot: Moonshot's list price; Z.ai: Z.ai's list price; Mistral: Mistral's price, served through OpenRouter). In llmwise you pay per message, not per token. Claude Haiku 5.5: the rate for prompts up to 100K tokens; $0.50 / $2.50 a million over that. Gemini 3.1 Pro (preview): the standard rate, for prompts up to 200K tokens. Gemini 3.8 Flash: an introductory price, through December 31, 2026. Grok 4.7: xAI charges more for very long prompts. Mistral Large 4: a sale price, half its list price of $1.36 / $4.18.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

How to choose

  1. Pick two or three real tasks from your own work, not toy examples.

  2. Run the first on an everyday model. If the answer is good enough, you're done: it didn't touch your monthly allowance.

  3. If it isn't, switch to a stronger model in the same chat and ask again. It sees the whole conversation.

  4. Keep the model that gets it right most often at a price you're happy with.

More in how to compare LLM models.

Bar chart: Prompts passed in our test runs, all 10 jobs. DeepSeek V4 Pro: 49 of 50; Gemini 3.1 Pro (preview): 49 of 50; Claude Opus 5.5: 49 of 50; GPT-6 Astra: 48 of 50; DeepSeek V4.1 Flash: 48 of 50; GLM 5.3: 48 of 50; GPT-6 Luna: 47 of 50; Grok 4.7: 47 of 50; Claude Sonnet 5.5: 47 of 50; Mistral Large 4: 47 of 50; GPT-6.1 Sol: 46 of 50; Claude Sonnet 5: 46 of 50; Gemini 3.8 Flash: 46 of 50; Kimi K3: 46 of 50; Claude Haiku 5.5: 45 of 50; GPT-6 Sol: 45 of 50; Claude Fable 5.1: 45 of 50; Claude Haiku 4.5: 43 of 50; GLM 5.3 Flash: 40 of 50.
Our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026: the same prompts for every model, each reply checked the same way.

Choosing an AI in 2026: questions

What is the best AI model in 2026?

There isn't one for everything, and any ranking is out of date by the next release. The practical answer is to match the model to the job: an everyday model for quick questions, a stronger one when the answer matters.

Does llmwise publish a ranking of AI models?

Yes: all 19 models ranked by our own test runs, overall, on the hard prompts alone and job by job, on our page on the best AI model right now. It's dated and changes when the results do, since rankings shift with every release; every prompt, reply and score is published.

Which models can I use for free?

Free is a one-time trial of 5 messages on every model but Claude Fable 5.1 and GPT-6 Astra (one of them can be on Claude Opus 5.5). It doesn't refill; a paid plan has everyday messages every day and a monthly allowance for the rest.

Pick per question, not per subscription

Start on an everyday model like GPT-6 Luna (60 a day on Pro) and move up when an answer isn't good enough: Claude Opus 5.5 gets up to 62 a month. The next model sees the whole chat.