GPT · Math
GPT-6 for math
llmwise has 4 of OpenAI's GPT models, from GPT-6 Luna (up to 60 messages a day on Pro) to GPT-6 Astra (up to 31 messages a month). We ran the same math prompts on every one and published every reply: which GPT model to use, from the results, what each costs per message, and how to get more out of it.
Based on 20 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
In our test runs on September 29, 2026, 3 models passed 5 of 5 math prompts, so these prompts don't pick one for hard problems. For value, GPT-6.1 Sol (5 of 5), 125 a month on Pro; for everyday math, GPT-6 Luna (4 of 5), from the daily count.
Our picks for math
Hard problems
Shared by 3 models
3 models passed 5 of 5, both hard ones: GPT-6 Astra, GPT-6.1 Sol and GPT-6 Sol. These prompts don't tell them apart, so they share the pick.
Best value
Passed 5 of 5 math prompts, with 125 a month on Pro.
Everyday
Passed 4 of 5 math prompts; an everyday model, so its messages come from the daily count (60 a day on Pro), not the monthly allowance.
These picks aren't our opinion: they're what the results below give, by these rules, among the GPT models in llmwise. They change when the results do.
- Hard problems: the model that passed the most prompts and, of those, the most hard ones. Models level on both share the pick: the prompts don't tell them apart, so we don't break the tie by price or by name.
- Best value: among the models that draw on the monthly allowance, the one with the most messages on Pro that passed no more than one prompt fewer than the top model. Ties go to the one that passed more, then to the lower cost per reply.
- Everyday: among the cheapest models on the page (the everyday models, which come from the daily count, when the page has any), the one that passed the most. Ties go to the one that passed more of the hard prompts, then to the lower cost per reply.
Our math test runs, model by model
How each model did on our 5 math prompts, what each reply counted as on Pro, and what it cost to run.
| Model | Passed | Hard ones | Messages used on Pro | Cost per reply | Time per reply |
|---|---|---|---|---|---|
| GPT-6 AstraOpenAI | 5 of 5 | 2 of 2 | 1 each, of 31 a month on Pro | $0.0085 | 2.7 s |
| GPT-6.1 SolOpenAI | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0009 | 2.0 s |
| GPT-6 SolOpenAI | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0018 | 2.2 s |
| GPT-6 LunaOpenAI | 4 of 5 | 2 of 2 | 1 each, of 60 a day on Pro | $0.0001 | 2.4 s |
One prompt, side by side
Each model's reply to “Four-digit numbers whose digits sum to 9”, as it came, with whether it passed and what it cost.
The prompt: Four-digit numbers whose digits sum to 9 (hard)
How many four-digit positive integers have digits that add up to 9? Show your working briefly, then end with a line that says "Final answer: " followed by the answer alone.
GPT-6 Astra
Passed: Final answer 165: right.
Let the digits be \(a,b,c,d\), where \(a\ge1\) and the others are nonnegative. We need \[ a+b+c+d=9. \] Set \(x=a-1\). Then \(x+b+c+d=8\). By stars and bars, the number of nonnegative solutions is \[ \binom{8+4-1}{4-1}=\binom{11}{3}=165. \] All solutions give valid digits, since \(a\le9\) and \(b,c,d\le8\). Final answer: 165378 tokens in, 169 out (35 of them reasoning) · 3.8 s · $0.0122 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·
GPT-6.1 Sol
Passed: Final answer 165: right.
Let the digits be \(a,b,c,d\), with \(a\ge1\). Set \(x=a-1\). Then \[ x+b+c+d=8, \] where all four variables are nonnegative. Each resulting digit is automatically at most 9. By stars and bars, the number of solutions is \[ \binom{8+4-1}{4-1}=\binom{11}{3}=165. \] Final answer: 165380 tokens in, 147 out (39 of them reasoning) · 2.7 s · $0.0011 · 1 message on Pro · answered by openai/gpt-6.1-sol via OpenAI ·
GPT-6 Luna
Passed: Final answer 165: right.
Let the first digit be at least 1. Subtract 1 from it, so the four digits become nonnegative integers summing to 8. By stars and bars, there are \(\binom{8+4-1}{4-1}=\binom{11}{3}=165\) possibilities. Final answer: 165378 tokens in, 127 out (50 of them reasoning) · 2.0 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·
The math prompts, and how they're scored
Final answer. Automatic. The reply's last “Final answer:” line must hold the right value.
Each prompt was sent the way llmwise sends a message in a side-by-side comparison, which offers no tools: the app's own system prompt, the model's own settings, and Pro's reply size limit (8,000 tokens). Read all five math prompts and how every reply was scored.
Prompts like these to try yourself
GPT in llmwise
| Model | On Pro | On Free | Context window | Images | PDFs | Reasoning | API price per 1M, in / out |
|---|---|---|---|---|---|---|---|
| GPT-6 AstraOpenAI | 31/mo on Pro | No | 1.05M tokens | Yes | Whole file | Yes | $10.00 / $50.00 |
| GPT-6.1 SolOpenAI | 125/mo on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes | $2.00 / $10.00 |
| GPT-6 SolOpenAI | 125/mo on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes | $2.00 / $10.00 |
| GPT-6 LunaOpenAI | 60/day on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes | $0.10 / $0.50 |
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
What matters for math
Step-by-step reasoning
Models that reason before answering work through the steps instead of jumping to a number, which matters most on multi-step problems.
Checking the arithmetic
Language models can slip on arithmetic. Running the calculation as code settles it.
Showing the work
A good answer shows its steps clearly enough for you to follow, and flags its assumptions.
GPT for math in llmwise
Readable formulas
Formulas in answers render as proper math, not raw LaTeX.
Run code (paid plans)
On a paid plan the model can run Python or Node.js in an isolated sandbox once you approve it, read the output and fix what failed. A run counts as 1 Claude Haiku 4.5 message and stops after 60 seconds.
Photos of problems
Attach a photo of a worksheet or a screenshot of a problem to a model that reads images.
The Tutor persona
Pick the Tutor persona to have concepts explained step by step, with checks that you followed.
In llmwise, a message to GPT goes to its maker, OpenAI, or through OpenRouter when llmwise can't reach the maker directly. See the Privacy Policy.
Tips
Ask for the steps, not just the answer.
Say what level you're at, so the explanation fits.
Ask the model to verify the final answer a second way.
For proofs, ask it to list its assumptions.
Questions
Is GPT good for math?
In our test runs on September 29, 2026, GPT-6 Astra passed 5 of 5, GPT-6.1 Sol passed 5 of 5, GPT-6 Sol passed 5 of 5, GPT-6 Luna passed 4 of 5 of our math prompts, against a best result of 5 of 5 among all 19 models. Every reply is published on this page and the methods page, so you can judge them yourself.
Which GPT model should I use for math?
Start with GPT-6 Luna for arithmetic, unit conversions and formulas you half remember, and move up to GPT-6 Astra for proofs, competition-style problems and anything where one slip ruins the answer.
Can I use GPT for math for free?
Yes: GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna are on the Free plan (a one-time trial of 5 messages). GPT-6 Astra needs a paid plan.
Can it check its own answer?
On a paid plan, ask the model to check its answer with code: it runs Python in a sandbox (after you approve) and compares. A run counts as 1 Claude Haiku 4.5 message.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.