Comparison · Math
ChatGPT vs Gemini for math
On llmwise Pro, GPT-6 Sol gets up to 125 messages a month and Gemini 3.1 Pro (preview) up to 125 messages a month. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself. We ran the same 5 math prompts on all 3 GPT models and all 2 Gemini models and published every reply: the results, then GPT-6 Sol against Gemini 3.8 Flash prompt by prompt, then how GPT and Gemini compare on price per message, context and files.
Based on 25 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
In our math test runs on September 27, 2026, GPT's 3 models passed 14 of 15; GPT-6 Astra and GPT-6 Sol each passed 5 of 5. Gemini's 2 models passed 10 of 10; Gemini 3.1 Pro and Gemini 3.8 Flash each passed 5 of 5. GPT-6 Sol and Gemini 3.8 Flash each passed 5 of the 5 prompts, so these math prompts don't split GPT and Gemini; the replies on this page show how they differ.
GPT and Gemini on our math test runs
Every GPT and Gemini model in llmwise on our 5 math prompts: how many replies passed, what each counted as on Pro, and what it cost to run.
| Model | Passed | Hard ones | Messages used on Pro | Cost per reply | Time per reply |
|---|---|---|---|---|---|
| GPT-6 AstraOpenAI | 5 of 5 | 2 of 2 | 1 each, of 31 a month on Pro | $0.0085 | 2.7 s |
| GPT-6 SolOpenAI | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0018 | 2.2 s |
| GPT-6 LunaOpenAI | 4 of 5 | 2 of 2 | 1 each, of 60 a day on Pro | $0.0001 | 2.4 s |
| Gemini 3.1 Pro (preview)Google | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0100 | 7.5 s |
| Gemini 3.8 FlashGoogle | 5 of 5 | 2 of 2 | 1 each, of 250 a month on Pro | $0.0014 | 3.8 s |
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
GPT vs Gemini, prompt by prompt
Each math prompt with GPT-6 Sol and Gemini 3.8 Flash side by side, as they replied: for each side, the model that passed the most math prompts (then the most hard ones, then the cheaper).
Math: GPT-6 Sol passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6 Sol answered 2.2× sooner at the median, 2.2 s against 4.7 s. Gemini 3.8 Flash cost 1.3× less, $0.0090 against $0.0071 for the 5 replies. Gemini 3.8 Flash's replies ran 128% longer, in tokens of reply, thinking not counted.
1. A discount, then sales tax
Everyday. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · Gemini: 2 of 2 models passed, Gemini 3.8 Flash the cheapest at $0.0007. The prompt in full.
GPT-6 Sol and Gemini 3.8 Flash both passed this one. GPT-6 Sol answered sooner (1.7 s to 2.5 s), and Gemini 3.8 Flash's reply cost less ($0.0007 to $0.0013); GPT-6 Sol wrote 45 tokens of reply, Gemini 3.8 Flash 89.
GPT-6 Sol
Passed: Final answer 66: right.
25% off $80 gives $80 × 0.75 = $60. Adding 10% sales tax gives $60 × 1.10 = $66. Final answer: $66
400 tokens in, 45 out · 1.7 s · $0.0013 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Gemini 3.8 Flash
Passed: Final answer 66: right.
1. Calculate the discount: $80 \times 0.25 = $20. 2. Calculate the discounted price: $80 - $20 = $60. 3. Calculate the sales tax: $60 \times 0.10 = $6. 4. Calculate the total cost: $60 + $6 = $66. Final answer: $66
430 tokens in, 89 out · 2.5 s · $0.0007 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·
2. Pens at 3 for $4
Everyday. GPT: 2 of 3 models passed, GPT-6 Sol the cheapest at $0.0024 · Gemini: 2 of 2 models passed, Gemini 3.8 Flash the cheapest at $0.0017. The prompt in full.
GPT-6 Sol and Gemini 3.8 Flash both passed this one. GPT-6 Sol answered sooner (3.3 s to 4.7 s), and Gemini 3.8 Flash's reply cost less ($0.0017 to $0.0024); GPT-6 Sol wrote 56 tokens of reply, Gemini 3.8 Flash 114.
GPT-6 Sol
Passed: Final answer 13.50: right.
Three groups of 3 pens cost \(3 \times \$4 = \$12\). One single pen costs \$1.50, bringing the total to \$13.50. Using fewer groups would cost more. Final answer: $13.50
402 tokens in, 164 out (108 of them reasoning) · 3.3 s · $0.0024 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Gemini 3.8 Flash
Passed: Final answer 3 sets of 3 pens and 1 single pen for 13.50: right.
To get exactly 10 pens at the lowest cost, buy as many 3-pen deals as possible: * 3 deals of 3 pens = 9 pens for $12.00 ($4.00 × 3) * 1 single pen = 1 pen for $1.50 Total cost: $12.00 + $1.50 = $13.50 Final answer: 3 sets of 3 pens and 1 single pen for $13.50
431 tokens in, 354 out (240 of them reasoning) · 4.7 s · $0.0017 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·
3. Compound interest over three years
Everyday. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · Gemini: 2 of 2 models passed, Gemini 3.8 Flash the cheapest at $0.0009. The prompt in full.
GPT-6 Sol and Gemini 3.8 Flash both passed this one. GPT-6 Sol answered sooner (1.4 s to 2.1 s), and Gemini 3.8 Flash's reply cost less ($0.0009 to $0.0014); GPT-6 Sol wrote 62 tokens of reply, Gemini 3.8 Flash 148.
GPT-6 Sol
Passed: Final answer 315.25: right.
After 3 years, the balance is \(2000(1.05)^3=\$2,315.25\). Subtracting the original deposit gives \(2315.25-2000=\$315.25\) in interest. Final answer: $315.25
402 tokens in, 62 out · 1.4 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Gemini 3.8 Flash
Passed: Final answer 315.25: right.
To find the total amount $A$ after 3 years: $A = P(1 + r)^t$ $A = 2000 \times (1 + 0.05)^3$ $A = 2000 \times (1.05)^3$ $A = 2000 \times 1.157625 = 2315.25$ To find the interest earned: $\text{Interest} = A - P = 2315.25 - 2000 = 315.25$ Final answer: $315.25429 tokens in, 148 out · 2.1 s · $0.0009 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·
4. Four-digit numbers whose digits sum to 9
Hard. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · Gemini: 2 of 2 models passed, Gemini 3.8 Flash the cheapest at $0.0022. The prompt in full.
GPT-6 Sol and Gemini 3.8 Flash both passed this one. GPT-6 Sol answered sooner (2.5 s to 4.7 s), and GPT-6 Sol's reply cost less ($0.0019 to $0.0022); GPT-6 Sol wrote 78 tokens of reply, Gemini 3.8 Flash 235.
GPT-6 Sol
Passed: Final answer 165: right.
Subtract 1 from the thousands digit. The four digits then become nonnegative integers that sum to 8. By stars and bars, the number of solutions is \(\binom{8+4-1}{4-1}=\binom{11}{3}=165\). Every solution gives a valid four-digit integer. Final answer: 165378 tokens in, 117 out (39 of them reasoning) · 2.5 s · $0.0019 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Gemini 3.8 Flash
Passed: Final answer 165: right.
Let the four digits be represented by $a, b, c,$ and $d$, where $a \in \{1, 2, \dots, 9\}$ and $b, c, d \in \{0, 1, \dots, 9\}$. We need to find the number of integer solutions to: $$a + b + c + d = 9$$ Let $a' = a - 1$ so that $a' \ge 0$. Substituting gives: $$a' + b + c + d = 8$$ Since the sum is $8$, no individual variable can exceed $8$, so the constraint that digits must be $\le 9$ is automatically satisfied. Using the stars and bars formula, the number of non-negative integer solutions is: $$\binom{8 + 4 - 1}{4 - 1} = \binom{11}{3} = \frac{11 \times 10 \times 9}{3 \times 2 \times 1} = 165$$ Final answer: 165404 tokens in, 495 out (260 of them reasoning) · 4.7 s · $0.0022 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·
5. The highest of three dice is a 5
Hard. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · Gemini: 2 of 2 models passed, Gemini 3.8 Flash the cheapest at $0.0018. The prompt in full.
GPT-6 Sol and Gemini 3.8 Flash both passed this one. GPT-6 Sol answered sooner (2.2 s to 4.9 s), and Gemini 3.8 Flash's reply cost less ($0.0018 to $0.0020); GPT-6 Sol wrote 91 tokens of reply, Gemini 3.8 Flash 172.
GPT-6 Sol
Passed: Final answer 61/216: right.
For the highest number to be exactly 5, all three dice must show 5 or less, but not all can show 4 or less. \[ P(\text{highest is }5)=\left(\frac56\right)^3-\left(\frac46\right)^3 =\frac{125-64}{216}=\frac{61}{216}. \] Final answer: 61/216393 tokens in, 122 out (31 of them reasoning) · 2.2 s · $0.0020 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Gemini 3.8 Flash
Passed: Final answer 61/216: right.
To find the probability that the highest number showing on three fair six-sided dice is exactly 5: 1. The total number of possible outcomes when rolling three dice is $6^3 = 216$. 2. The number of outcomes where all three dice show a number less than or equal to 5 is $5^3 = 125$. 3. The number of outcomes where all three dice show a number less than or equal to 4 (meaning no 5 is rolled) is $4^3 = 64$. 4. The number of outcomes where the maximum number is exactly 5 is $125 - 64 = 61$. Thus, the probability is $\frac{61}{216}$. Final answer: 61/216421 tokens in, 392 out (220 of them reasoning) · 4.9 s · $0.0018 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·
Each GPT model against each Gemini model
Every GPT model against every Gemini model on the same 5 math prompts: passes, the median wait and what the replies cost. A lead under 10% counts as close.
GPT-6 Astra vs Gemini 3.1 Pro: 5 of 5 each; GPT-6 Astra answered 3.2× sooner at the median and cost 1.2× less.
GPT-6 Astra vs Gemini 3.8 Flash: 5 of 5 each; GPT-6 Astra answered 1.8× sooner at the median and Gemini 3.8 Flash cost 6.0× less.
GPT-6 Sol vs Gemini 3.1 Pro: 5 of 5 each; GPT-6 Sol answered 3.8× sooner at the median and cost 5.5× less. GPT-6 Sol vs Gemini 3.1 Pro (preview), on every job.
GPT-6 Sol vs Gemini 3.8 Flash: 5 of 5 each; GPT-6 Sol answered 2.2× sooner at the median and Gemini 3.8 Flash cost 1.3× less. GPT-6 Sol vs Gemini 3.8 Flash, on every job.
GPT-6 Luna vs Gemini 3.1 Pro: GPT-6 Luna 4 of 5, Gemini 3.1 Pro 5; GPT-6 Luna answered 3.8× sooner at the median and cost 101.7× less.
GPT-6 Luna vs Gemini 3.8 Flash: GPT-6 Luna 4 of 5, Gemini 3.8 Flash 5; GPT-6 Luna answered 2.1× sooner at the median and cost 14.5× less. GPT-6 Luna vs Gemini 3.8 Flash, on every job.
How the math replies are scored
Every GPT and Gemini reply above was checked the same way as every other model's, by the rules published with the prompts: how the math prompts are scored, and each one in full.
The differences at a glance
What follows from each model's facts in our catalog.
The lineups
GPT: 3 models, GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. Gemini: 2 models, Gemini 3.1 Pro (preview) and Gemini 3.8 Flash.
Price per message
The least expensive GPT model is GPT-6 Luna (60 messages a day on Pro); the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro).
Context window
GPT goes up to 1.05M tokens (GPT-6 Astra); Gemini up to 1.05M tokens (Gemini 3.1 Pro (preview)).
Images and PDFs
Every model here reads images. Every model here takes a PDF as a whole file.
On the Free plan
Free's one-time trial of 5 messages covers GPT-6 Sol, GPT-6 Luna, Gemini 3.1 Pro (preview), and Gemini 3.8 Flash. Paid plans have every model, with messages every month.
Every GPT and Gemini model's context window, files and API price: ChatGPT vs Gemini.
Where your messages go
In llmwise, a message to GPT or Gemini goes to the model's maker, or through OpenRouter when llmwise can't reach the maker directly. The Privacy Policy has the details.
Questions
Which is better, ChatGPT or Gemini for math?
In our math test runs on September 27, 2026, GPT's 3 models passed 14 of 15; GPT-6 Astra and GPT-6 Sol each passed 5 of 5. Gemini's 2 models passed 10 of 10; Gemini 3.1 Pro and Gemini 3.8 Flash each passed 5 of 5. GPT-6 Sol and Gemini 3.8 Flash each passed 5 of the 5 prompts, so these math prompts don't split GPT and Gemini; the replies on this page show how they differ.
Is GPT or Gemini cheaper?
In llmwise, the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro), and the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro). At API list prices (September 2026), a typical message of 4,000 tokens in and 700 out costs $0.0008 on GPT-6 Luna and $0.0056 on Gemini 3.8 Flash.
Can I use GPT and Gemini in the same chat?
Yes. Ask GPT-6 Sol a question, then switch the picker to Gemini 3.8 Flash and ask again: Gemini 3.8 Flash sees the whole conversation, GPT-6 Sol's answer included.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.