Skip to content

Comparison · 3 families

ChatGPT vs Claude vs Gemini

On llmwise Pro, GPT-6 Sol up to 125 messages a month, Claude Sonnet 5.5 up to 125, and Gemini 3.1 Pro (preview) up to 125. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself. We ran the same 50 prompts across 10 jobs on every GPT, Claude, and Gemini model and published every reply: here's each family job by job, then each one's own subscription and every model's price.

Based on 500 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

In our test runs on September 28, 2026, the same 50 prompts across 10 jobs: GPT's 3 models passed 140 of 150 replies, Claude's 5 models passed 230 of 250 replies, and Gemini's 2 models passed 95 of 100 replies. Job by job, all 3 families' best models shared the top result on 8 of the 10 jobs; GPT's alone had it on writing; GPT's trailed on customer support; Claude's trailed on writing; Gemini's trailed on writing.

GPT vs Claude vs Gemini, job by job

Each family's best model on each job, by our published rule (the most prompts passed, then the most hard ones, then the cheaper), and what it passed. The top result on each job is marked.

GPT, Claude, Gemini: each family's best result on each job
JobGPTClaudeGemini
Coding5 of 5 · top3 models tied5 of 5 · top4 models tied5 of 5 · top2 models tied
Writing5 of 5 · top2 models tied4 of 5Claude Sonnet 54 of 5Gemini 3.1 Pro
Math5 of 5 · top2 models tied5 of 5 · top5 models tied5 of 5 · top2 models tied
Summarization5 of 5 · topGPT-6 Astra5 of 5 · top2 models tied5 of 5 · top2 models tied
Data analysis5 of 5 · top3 models tied5 of 5 · top3 models tied5 of 5 · topGemini 3.1 Pro
Customer support4 of 5GPT-6 Luna5 of 5 · top2 models tied5 of 5 · topGemini 3.1 Pro
Translation5 of 5 · top2 models tied5 of 5 · top5 models tied5 of 5 · top2 models tied
SQL5 of 5 · top3 models tied5 of 5 · top5 models tied5 of 5 · top2 models tied
RAG and answering from documents5 of 5 · top3 models tied5 of 5 · top5 models tied5 of 5 · top2 models tied
Agents and tool use5 of 5 · top3 models tied5 of 5 · top5 models tied5 of 5 · top2 models tied
Each job has 5 prompts; “top” marks the family (or families) whose best model passed the most of them. A job with a head-to-head page links to it.

Every model, every job

All 10 GPT, Claude, and Gemini models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.

Every GPT, Claude, and Gemini model across our test runs
ModelPassedHard onesCost per replyOn Pro
GPT-6 AstraOpenAI48 of 5020 of 20$0.0117Up to 31 a month
GPT-6 SolOpenAI45 of 5019 of 20$0.0026Up to 125 a month
GPT-6 LunaOpenAI47 of 5020 of 20$0.00012Up to 60 a day
Claude Fable 5.1Anthropic45 of 5017 of 20$0.0200Up to 31 a month
Claude Opus 5.5Anthropic49 of 5019 of 20$0.0107Up to 62 a month
Claude Sonnet 5.5Anthropic47 of 5019 of 20$0.0038Up to 125 a month
Claude Sonnet 5Anthropic46 of 5019 of 20$0.0041Up to 125 a month
Claude Haiku 4.5Anthropic43 of 5015 of 20$0.0015Up to 250 a month
Gemini 3.1 Pro (preview)Google49 of 5019 of 20$0.0107Up to 125 a month
Gemini 3.8 FlashGoogle46 of 5018 of 20$0.0013Up to 250 a month

Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.

One writing prompt, 3 replies

GPT-6 Luna, Claude Sonnet 5, and Gemini 3.1 Pro, each its family's pick on writing by the rule above, on the same prompt: every reply as it came.

A hard prompt: “A product announcement with five rules”, in full.

  • GPT-6 Luna

    Passed: Graded 4.0 of 5 on average (lowest 3).

    Meet the Nordvik 2, a desk lamp designed to bring flexible light to your workspace. Choose from five brightness levels, charge it via USB-C, and turn its arm up to 270 degrees to direct light where you need it. The Nordvik 2 costs €79 and ships on 1 December. Ready to brighten your desk?

    406 tokens in, 73 out · 1.5 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  • Claude Sonnet 5

    Passed: Graded 4.3 of 5 on average (lowest 4).

    **Introducing the Nordvik 2 Desk Lamp**
    
    Meet the Nordvik 2, thoughtfully redesigned for how you actually work. Five-step dimming lets you fine-tune brightness for any task, from focused reading to ambient evening light. The 270-degree rotating arm adjusts to virtually any angle, giving you precise control over where light falls. USB-C charging means one cable does it all, keeping your desk clutter-free and compatible with everything else you own.
    
    Priced at €79, the Nordvik 2 brings premium design and functionality within easy reach. It begins shipping on 1 December, just in time to brighten your winter workspace.
    
    Ready to see how the right light changes the way you work?

    554 tokens in, 235 out · 4.5 s · $0.0035 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

  • Gemini 3.1 Pro

    Passed: Graded 4.0 of 5 on average (lowest 3).

    Introducing the Nordvik 2 desk lamp, designed to illuminate your workspace with precision. This new model features a highly flexible arm that turns 270 degrees, putting light exactly where you need it. You can easily tailor the brightness to suit any task, as it dims in five steps. Built for modern convenience, it also charges from USB-C to keep your setup clean and wire-free. The Nordvik 2 costs €79 and ships on 1 December. Are you ready to upgrade your daily lighting?

    435 tokens in, 715 out (609 of them reasoning) · 9.2 s · $0.0095 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

Their own subscriptions

Each company's own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.

llmwise Pro, $20 a month, has all 10 of these models in one chat, from one monthly allowance: GPT-6 Sol up to 125 messages a month, Claude Sonnet 5.5 up to 125, and Gemini 3.1 Pro (preview) up to 125.

Every GPT, Claude, and Gemini model in llmwise

GPT, Claude, and Gemini models compared
ModelOn ProOn FreeContext windowImagesPDFsReasoningAPI price per 1M, in / out
GPT-6 AstraOpenAI31/mo on ProNo1.05M tokensYesWhole fileYes$10.00 / $50.00
GPT-6 SolOpenAI125/mo on ProYes1.05M tokensYesWhole fileYes$2.00 / $10.00
GPT-6 LunaOpenAI60/day on ProYes1.05M tokensYesWhole fileYes$0.10 / $0.50
Claude Fable 5.1Anthropic31/mo on ProNo1M tokensYesWhole fileYes$10.00 / $50.00
Claude Opus 5.5Anthropic62/mo on ProNo1M tokensYesWhole fileYes$4.00 / $20.00
Claude Sonnet 5.5Anthropic125/mo on ProYes1M tokensYesWhole fileYes$2.00 / $10.00
Claude Sonnet 5Anthropic125/mo on ProYes1M tokensYesWhole fileYes$2.00 / $10.00
Claude Haiku 4.5Anthropic250/mo on ProYes200K tokensYesWhole fileNo$1.00 / $5.00
Gemini 3.1 Pro (preview)Google125/mo on ProYes1.05M tokensYesWhole fileYes$2.00 / $12.00
Gemini 3.8 FlashGoogle250/mo on ProYes1.05M tokensYesWhole fileYes$0.75 / $3.75
Each badge is how many messages Pro gets on the model: a month’s, or a day’s on an everyday model. Free is a one-time trial of 5 messages on the models marked. “Text only” models get the text of a PDF, not the file. API prices are the per-token prices in our model catalog as of September 2026 (OpenAI: OpenAI's list price; Anthropic: Anthropic's list price; Google: Google's list price). In llmwise you pay per message, not per token. Gemini 3.1 Pro (preview): the standard rate, for prompts up to 200K tokens. Gemini 3.8 Flash: an introductory price, through December 31, 2026.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

More head-to-heads

These families two at a time, and the other pages about several.

Try GPT, Claude, and Gemini on your own question

  1. Start a chat on GPT and ask a real question from your work, with any files it needs.

  2. Switch the model picker to Claude, then Gemini in turn, and ask each for its answer. Each sees the whole chat, the other answers included.

  3. Compare them on what matters to you: accuracy, tone, length, how much you'd have to fix.

  4. The badge by each model in the picker shows what a message uses before you send it.

Where your messages go

In llmwise, a message to GPT, Claude, or Gemini goes to the model's maker, or through OpenRouter when llmwise can't reach the maker directly. The Privacy Policy has the details.

Questions

Which is better, ChatGPT, Claude, or Gemini?

In our test runs on September 28, 2026, the same 50 prompts across 10 jobs: GPT's 3 models passed 140 of 150 replies, Claude's 5 models passed 230 of 250 replies, and Gemini's 2 models passed 95 of 100 replies. Job by job, all 3 families' best models shared the top result on 8 of the 10 jobs; GPT's alone had it on writing; GPT's trailed on customer support; Claude's trailed on writing; Gemini's trailed on writing.

Which is cheaper, GPT, Claude, or Gemini?

At API list prices (September 2026), the least expensive model of each family for a typical message is GPT-6 Luna ($0.0008), Claude Haiku 4.5 ($0.0075), and Gemini 3.8 Flash ($0.0056). In llmwise you pay per message, and every one of them is on the same plan.

Can I use GPT, Claude, and Gemini in the same chat?

Yes. Pick GPT, Claude, and Gemini models message by message in one chat; each model sees the whole conversation, the others' answers included.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.