Comparison · Math
Claude vs ChatGPT for math
On llmwise Pro, Claude Sonnet 5.5 gets up to 125 messages a month and GPT-6 Sol up to 125 messages a month. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself. We ran the same 5 math prompts on all 5 Claude models and all 3 GPT models and published every reply: the results, then Claude Haiku 4.5 against GPT-6 Sol prompt by prompt, then how Claude and GPT compare on price per message, context and files.
Based on 40 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
In our math test runs on September 28, 2026, Claude's 5 models passed 25 of 25; Claude Fable 5.1, Claude Opus 5.5 and 3 more each passed 5 of 5. GPT's 3 models passed 14 of 15; GPT-6 Astra and GPT-6 Sol each passed 5 of 5. Claude Haiku 4.5 and GPT-6 Sol each passed 5 of the 5 prompts, so these math prompts don't split Claude and GPT; the replies on this page show how they differ.
Claude and GPT on our math test runs
Every Claude and GPT model in llmwise on our 5 math prompts: how many replies passed, what each counted as on Pro, and what it cost to run.
| Model | Passed | Hard ones | Messages used on Pro | Cost per reply | Time per reply |
|---|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 5 of 5 | 2 of 2 | 1 each, of 31 a month on Pro | $0.0116 | 4.6 s |
| Claude Opus 5.5Anthropic | 5 of 5 | 2 of 2 | 1 each, of 62 a month on Pro | $0.0060 | 4.1 s |
| Claude Sonnet 5.5Anthropic | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0025 | 1.7 s |
| Claude Sonnet 5Anthropic | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0033 | 4.1 s |
| Claude Haiku 4.5Anthropic | 5 of 5 | 2 of 2 | 1 each, of 250 a month on Pro | $0.0015 | 2.4 s |
| GPT-6 AstraOpenAI | 5 of 5 | 2 of 2 | 1 each, of 31 a month on Pro | $0.0085 | 2.7 s |
| GPT-6 SolOpenAI | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0018 | 2.2 s |
| GPT-6 LunaOpenAI | 4 of 5 | 2 of 2 | 1 each, of 60 a day on Pro | $0.0001 | 2.4 s |
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
Claude vs GPT, prompt by prompt
Each math prompt with Claude Haiku 4.5 and GPT-6 Sol side by side, as they replied: for each side, the model that passed the most math prompts (then the most hard ones, then the cheaper).
Math: Claude Haiku 4.5 passed 5 of 5 and GPT-6 Sol 5 of 5. Their median waits were close, 2.2 s against 2.2 s. Claude Haiku 4.5 cost 1.2× less, $0.0077 against $0.0090 for the 5 replies. Claude Haiku 4.5's replies ran 236% longer, in tokens of reply, thinking not counted.
1. A discount, then sales tax
Everyday. Claude: 5 of 5 models passed, Claude Haiku 4.5 the cheapest at $0.0007 · GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001. The prompt in full.
Claude Haiku 4.5 and GPT-6 Sol both passed this one. Claude Haiku 4.5 answered sooner (1.1 s to 1.7 s), and Claude Haiku 4.5's reply cost less ($0.0007 to $0.0013); Claude Haiku 4.5 wrote 63 tokens of reply, GPT-6 Sol 45.
Claude Haiku 4.5
Passed: Final answer 66: right.
**Working:** 1. Original price: $80 2. Apply 25% discount: $80 × 0.75 = $60 3. Add 10% sales tax: $60 × 1.10 = $66 Final answer: $66
429 tokens in, 63 out · 1.1 s · $0.0007 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·
GPT-6 Sol
Passed: Final answer 66: right.
25% off $80 gives $80 × 0.75 = $60. Adding 10% sales tax gives $60 × 1.10 = $66. Final answer: $66
400 tokens in, 45 out · 1.7 s · $0.0013 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
2. Pens at 3 for $4
Everyday. Claude: 5 of 5 models passed, Claude Haiku 4.5 the cheapest at $0.0014 · GPT: 2 of 3 models passed, GPT-6 Sol the cheapest at $0.0024. The prompt in full.
Claude Haiku 4.5 and GPT-6 Sol both passed this one. Claude Haiku 4.5 answered sooner (2.2 s to 3.3 s), and Claude Haiku 4.5's reply cost less ($0.0014 to $0.0024); Claude Haiku 4.5 wrote 199 tokens of reply, GPT-6 Sol 56.
Claude Haiku 4.5
Passed: Final answer Buy 3 packs of 3 pens and 1 single pen for 13.50: right.
# Pen Purchase Problem **Option 1: Buy 3-packs only** - 10 pens ÷ 3 = 3 packs with 1 pen left over - 3 packs × $4 = $12 - 1 single pen × $1.50 = $1.50 - Total: $13.50 **Option 2: Buy all single pens** - 10 pens × $1.50 = $15.00 **Option 3: Buy 3 packs + adjust** - Same as Option 1: $13.50 The cheapest way is to buy three 3-packs ($12) and one single pen ($1.50). Final answer: Buy 3 packs of 3 pens and 1 single pen for $13.50
434 tokens in, 199 out · 2.2 s · $0.0014 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·
GPT-6 Sol
Passed: Final answer 13.50: right.
Three groups of 3 pens cost \(3 \times \$4 = \$12\). One single pen costs \$1.50, bringing the total to \$13.50. Using fewer groups would cost more. Final answer: $13.50
402 tokens in, 164 out (108 of them reasoning) · 3.3 s · $0.0024 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
3. Compound interest over three years
Everyday. Claude: 5 of 5 models passed, Claude Haiku 4.5 the cheapest at $0.0013 · GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001. The prompt in full.
Claude Haiku 4.5 and GPT-6 Sol both passed this one. GPT-6 Sol answered sooner (1.4 s to 3.5 s), and cost about the same; Claude Haiku 4.5 wrote 173 tokens of reply, GPT-6 Sol 62.
Claude Haiku 4.5
Passed: Final answer 315.25: right.
I'll calculate the compound interest for you. **Formula:** A = P(1 + r)^n Where: - P = Principal = $2,000 - r = Annual interest rate = 0.05 - n = Number of years = 3 **Calculation:** - Year 1: $2,000 × 1.05 = $2,100.00 - Year 2: $2,100 × 1.05 = $2,205.00 - Year 3: $2,205 × 1.05 = $2,315.25 **Interest earned:** $2,315.25 - $2,000.00 = $315.25 Final answer: $315.25
430 tokens in, 173 out · 3.5 s · $0.0013 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·
GPT-6 Sol
Passed: Final answer 315.25: right.
After 3 years, the balance is \(2000(1.05)^3=\$2,315.25\). Subtracting the original deposit gives \(2315.25-2000=\$315.25\) in interest. Final answer: $315.25
402 tokens in, 62 out · 1.4 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
4. Four-digit numbers whose digits sum to 9
Hard. Claude: 5 of 5 models passed, Claude Haiku 4.5 the cheapest at $0.0024 · GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001. The prompt in full.
Claude Haiku 4.5 and GPT-6 Sol both passed this one. GPT-6 Sol answered sooner (2.5 s to 2.9 s), and GPT-6 Sol's reply cost less ($0.0019 to $0.0024); Claude Haiku 4.5 wrote 404 tokens of reply, GPT-6 Sol 78.
Claude Haiku 4.5
Passed: Final answer 165: right.
I need to find four-digit positive integers where the digits sum to 9. A four-digit positive integer has the form $\overline{abcd}$ where: - $a \in \{1,2,...,9\}$ (first digit can't be 0) - $b,c,d \in \{0,1,...,9\}$ - $a + b + c + d = 9$ **Method: Change of variables** Let $a' = a - 1$, so $a' \in \{0,1,...,8\}$. Then: $a' + 1 + b + c + d = 9$ This gives: $a' + b + c + d = 8$ Now I need to count non-negative integer solutions where: - $a' \in \{0,1,...,8\}$ - $b,c,d \in \{0,1,...,9\}$ - $a' + b + c + d = 8$ Since the sum is only 8, each variable is automatically ≤ 8 (and thus $a' \leq 8$ and $b,c,d \leq 9$ are automatically satisfied). This is a standard stars and bars problem: distribute 8 identical objects into 4 distinct bins with no restrictions. The number of solutions is: $$\binom{8 + 4 - 1}{4 - 1} = \binom{11}{3} = \frac{11 \times 10 \times 9}{3 \times 2 \times 1} = \frac{990}{6} = 165$$ Final answer: 165404 tokens in, 404 out · 2.9 s · $0.0024 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·
GPT-6 Sol
Passed: Final answer 165: right.
Subtract 1 from the thousands digit. The four digits then become nonnegative integers that sum to 8. By stars and bars, the number of solutions is \(\binom{8+4-1}{4-1}=\binom{11}{3}=165\). Every solution gives a valid four-digit integer. Final answer: 165378 tokens in, 117 out (39 of them reasoning) · 2.5 s · $0.0019 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
5. The highest of three dice is a 5
Hard. Claude: 5 of 5 models passed, Claude Haiku 4.5 the cheapest at $0.0018 · GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001. The prompt in full.
Claude Haiku 4.5 and GPT-6 Sol both passed this one. They took about as long (Claude Haiku 4.5 2.2 s, GPT-6 Sol 2.2 s), and Claude Haiku 4.5's reply cost less ($0.0018 to $0.0020); Claude Haiku 4.5 wrote 276 tokens of reply, GPT-6 Sol 91.
Claude Haiku 4.5
Passed: Final answer 61/216: right.
I need to find the probability that the maximum of three dice is exactly 5. **Working:** For the maximum to be exactly 5: - At least one die shows 5 - No die shows 6 **Count favorable outcomes:** - All three dice must show values from {1, 2, 3, 4, 5} - At least one die must show 5 Using complementary counting: - Total ways to roll three dice with values in {1,2,3,4,5}: 5³ = 125 - Ways with no 5 (all from {1,2,3,4}): 4³ = 64 - Ways with at least one 5: 125 - 64 = 61 **Total possible outcomes:** 6³ = 216 **Probability:** 61/216 Let me verify this is in lowest terms by checking if 61 and 216 share factors: - 61 is prime - 216 = 2³ × 3³ - Since 61 is prime and doesn't equal 2 or 3, gcd(61, 216) = 1 Final answer: 61/216420 tokens in, 276 out · 2.2 s · $0.0018 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·
GPT-6 Sol
Passed: Final answer 61/216: right.
For the highest number to be exactly 5, all three dice must show 5 or less, but not all can show 4 or less. \[ P(\text{highest is }5)=\left(\frac56\right)^3-\left(\frac46\right)^3 =\frac{125-64}{216}=\frac{61}{216}. \] Final answer: 61/216393 tokens in, 122 out (31 of them reasoning) · 2.2 s · $0.0020 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Each Claude model against each GPT model
Every Claude model against every GPT model on the same 5 math prompts: passes, the median wait and what the replies cost. A lead under 10% counts as close.
Claude Fable 5.1 vs GPT-6 Astra: 5 of 5 each; GPT-6 Astra answered 1.8× sooner at the median and cost 1.4× less. Claude Fable 5.1 vs GPT-6 Astra, on every job.
Claude Fable 5.1 vs GPT-6 Sol: 5 of 5 each; GPT-6 Sol answered 2.1× sooner at the median and cost 6.4× less.
Claude Fable 5.1 vs GPT-6 Luna: Claude Fable 5.1 5 of 5, GPT-6 Luna 4; GPT-6 Luna answered 2.1× sooner at the median and cost 118.2× less.
Claude Opus 5.5 vs GPT-6 Astra: 5 of 5 each; GPT-6 Astra answered 1.5× sooner at the median and Claude Opus 5.5 cost 1.4× less. Claude Opus 5.5 vs GPT-6 Astra, on every job.
Claude Opus 5.5 vs GPT-6 Sol: 5 of 5 each; GPT-6 Sol answered 1.8× sooner at the median and cost 3.3× less. Claude Opus 5.5 vs GPT-6 Sol, on every job.
Claude Opus 5.5 vs GPT-6 Luna: Claude Opus 5.5 5 of 5, GPT-6 Luna 4; GPT-6 Luna answered 1.8× sooner at the median and cost 60.5× less.
Claude Sonnet 5.5 vs GPT-6 Astra: 5 of 5 each; Claude Sonnet 5.5 answered 1.6× sooner at the median and cost 3.5× less.
Claude Sonnet 5.5 vs GPT-6 Sol: 5 of 5 each; Claude Sonnet 5.5 answered 1.3× sooner at the median and GPT-6 Sol cost 1.4× less. Claude Sonnet 5.5 vs GPT-6 Sol, on every job.
Claude Sonnet 5.5 vs GPT-6 Luna: Claude Sonnet 5.5 5 of 5, GPT-6 Luna 4; Claude Sonnet 5.5 answered 1.3× sooner at the median and GPT-6 Luna cost 25.1× less.
Claude Sonnet 5 vs GPT-6 Astra: 5 of 5 each; GPT-6 Astra answered 1.5× sooner at the median and Claude Sonnet 5 cost 2.6× less.
Claude Sonnet 5 vs GPT-6 Sol: 5 of 5 each; GPT-6 Sol answered 1.8× sooner at the median and cost 1.8× less. Claude Sonnet 5 vs GPT-6 Sol, on every job.
Claude Sonnet 5 vs GPT-6 Luna: Claude Sonnet 5 5 of 5, GPT-6 Luna 4; GPT-6 Luna answered 1.8× sooner at the median and cost 33.5× less.
Claude Haiku 4.5 vs GPT-6 Astra: 5 of 5 each; Claude Haiku 4.5 answered 1.2× sooner at the median and cost 5.6× less.
Claude Haiku 4.5 vs GPT-6 Sol: 5 of 5 each; they took about as long and Claude Haiku 4.5 cost 1.2× less. Claude Haiku 4.5 vs GPT-6 Sol, on every job.
Claude Haiku 4.5 vs GPT-6 Luna: Claude Haiku 4.5 5 of 5, GPT-6 Luna 4; they took about as long and GPT-6 Luna cost 15.6× less. Claude Haiku 4.5 vs GPT-6 Luna, on every job.
How the math replies are scored
Every Claude and GPT reply above was checked the same way as every other model's, by the rules published with the prompts: how the math prompts are scored, and each one in full.
The differences at a glance
What follows from each model's facts in our catalog.
The lineups
Claude: 5 models, Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, Claude Sonnet 5, and Claude Haiku 4.5. GPT: 3 models, GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna.
Price per message
The least expensive Claude model is Claude Haiku 4.5 (250 messages a month on Pro); the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro).
Context window
Claude goes up to 1M tokens (Claude Fable 5.1); GPT up to 1.05M tokens (GPT-6 Astra).
Images and PDFs
Every model here reads images. Every model here takes a PDF as a whole file.
On the Free plan
Free's one-time trial of 5 messages covers Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 4.5, GPT-6 Sol, and GPT-6 Luna. Paid plans have every model, with messages every month.
Every Claude and GPT model's context window, files and API price: Claude vs ChatGPT.
Where your messages go
In llmwise, a message to Claude or GPT goes to the model's maker, or through OpenRouter when llmwise can't reach the maker directly. The Privacy Policy has the details.
Questions
Which is better, Claude or ChatGPT for math?
In our math test runs on September 28, 2026, Claude's 5 models passed 25 of 25; Claude Fable 5.1, Claude Opus 5.5 and 3 more each passed 5 of 5. GPT's 3 models passed 14 of 15; GPT-6 Astra and GPT-6 Sol each passed 5 of 5. Claude Haiku 4.5 and GPT-6 Sol each passed 5 of the 5 prompts, so these math prompts don't split Claude and GPT; the replies on this page show how they differ.
Is Claude or GPT cheaper?
In llmwise, the least expensive Claude model is Claude Haiku 4.5 (250 messages a month on Pro), and the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro). At API list prices (September 2026), a typical message of 4,000 tokens in and 700 out costs $0.0075 on Claude Haiku 4.5 and $0.0008 on GPT-6 Luna.
Can I use Claude and GPT in the same chat?
Yes. Ask Claude Haiku 4.5 a question, then switch the picker to GPT-6 Sol and ask again: GPT-6 Sol sees the whole conversation, Claude Haiku 4.5's answer included.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.