Comparison · Math
ChatGPT vs Grok for math
On llmwise Pro, GPT-6 Sol gets up to 125 messages a month and Grok 4.7 up to 250 messages a month. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself. We ran the same 5 math prompts on all 3 GPT models and Grok's one model and published every reply: the results, then GPT-6 Sol against Grok 4.7 prompt by prompt, then how GPT and Grok compare on price per message, context and files.
Based on 20 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
In our math test runs on September 27, 2026, GPT's 3 models passed 14 of 15; GPT-6 Astra and GPT-6 Sol each passed 5 of 5. Grok's one model passed 5 of 5 ($0.0050 a reply). GPT-6 Sol and Grok 4.7 each passed 5 of the 5 prompts, so these math prompts don't split GPT and Grok; the replies on this page show how they differ.
GPT and Grok on our math test runs
Every GPT and Grok model in llmwise on our 5 math prompts: how many replies passed, what each counted as on Pro, and what it cost to run.
| Model | Passed | Hard ones | Messages used on Pro | Cost per reply | Time per reply |
|---|---|---|---|---|---|
| GPT-6 AstraOpenAI | 5 of 5 | 2 of 2 | 1 each, of 31 a month on Pro | $0.0085 | 2.7 s |
| GPT-6 SolOpenAI | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0018 | 2.2 s |
| GPT-6 LunaOpenAI | 4 of 5 | 2 of 2 | 1 each, of 60 a day on Pro | $0.0001 | 2.4 s |
| Grok 4.7xAI | 5 of 5 | 2 of 2 | 1 each, of 250 a month on Pro | $0.0050 | 14.0 s |
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
GPT vs Grok, prompt by prompt
Each math prompt with GPT-6 Sol and Grok 4.7 side by side, as they replied: for each side, the model that passed the most math prompts (then the most hard ones, then the cheaper).
Math: GPT-6 Sol passed 5 of 5 and Grok 4.7 5 of 5. GPT-6 Sol answered 4.1× sooner at the median, 2.2 s against 9.1 s. GPT-6 Sol cost 2.8× less, $0.0090 against $0.0251 for the 5 replies. Grok 4.7's replies ran 62% longer, in tokens of reply, thinking not counted.
1. A discount, then sales tax
Everyday. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · Grok: 1 of 1 model passed, Grok 4.7 at $0.0027. The prompt in full.
GPT-6 Sol and Grok 4.7 both passed this one. GPT-6 Sol answered sooner (1.7 s to 3.9 s), and GPT-6 Sol's reply cost less ($0.0013 to $0.0027); GPT-6 Sol wrote 45 tokens of reply, Grok 4.7 40.
GPT-6 Sol
Passed: Final answer 66: right.
25% off $80 gives $80 × 0.75 = $60. Adding 10% sales tax gives $60 × 1.10 = $66. Final answer: $66
400 tokens in, 45 out · 1.7 s · $0.0013 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Grok 4.7
Passed: Final answer 66: right.
Discounted price: \(80 \times 0.75 = 60\). With tax: \(60 \times 1.10 = 66\). Final answer: $66
1,634 tokens in, 307 out (267 of them reasoning) · 3.9 s · $0.0027 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
2. Pens at 3 for $4
Everyday. GPT: 2 of 3 models passed, GPT-6 Sol the cheapest at $0.0024 · Grok: 1 of 1 model passed, Grok 4.7 at $0.0087. The prompt in full.
GPT-6 Sol and Grok 4.7 both passed this one. GPT-6 Sol answered sooner (3.3 s to 40.7 s), and GPT-6 Sol's reply cost less ($0.0024 to $0.0087); GPT-6 Sol wrote 56 tokens of reply, Grok 4.7 96.
GPT-6 Sol
Passed: Final answer 13.50: right.
Three groups of 3 pens cost \(3 \times \$4 = \$12\). One single pen costs \$1.50, bringing the total to \$13.50. Using fewer groups would cost more. Final answer: $13.50
402 tokens in, 164 out (108 of them reasoning) · 3.3 s · $0.0024 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Grok 4.7
Passed: Final answer 3 bundles + 1 single ( 13.50): right.
Buy as many 3-for-$4 bundles as possible, since each bundle pen costs about $1.33 versus $1.50 for a single. For exactly 10 pens the options are 0–3 bundles plus singles; the cost formula \(15 - 0.5k\) (where \(k\) is the number of bundles) is minimized at \(k=3\). Final answer: 3 bundles + 1 single ($13.50)
1,636 tokens in, 1,547 out (1,451 of them reasoning) · 40.7 s · $0.0087 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
3. Compound interest over three years
Everyday. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · Grok: 1 of 1 model passed, Grok 4.7 at $0.0031. The prompt in full.
GPT-6 Sol and Grok 4.7 both passed this one. GPT-6 Sol answered sooner (1.4 s to 5.5 s), and GPT-6 Sol's reply cost less ($0.0014 to $0.0031); GPT-6 Sol wrote 62 tokens of reply, Grok 4.7 91.
GPT-6 Sol
Passed: Final answer 315.25: right.
After 3 years, the balance is \(2000(1.05)^3=\$2,315.25\). Subtracting the original deposit gives \(2315.25-2000=\$315.25\) in interest. Final answer: $315.25
402 tokens in, 62 out · 1.4 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Grok 4.7
Passed: Final answer 315.25: right.
Principal \(P = 2000\), rate \(r = 0.05\), time \(t = 3\). \[ A = 2000 \times (1.05)^3 = 2000 \times 1.157625 = 2315.25 \] Interest earned \(= 2315.25 - 2000 = 315.25\). Final answer: $315.25
1,636 tokens in, 460 out (369 of them reasoning) · 5.5 s · $0.0031 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
4. Four-digit numbers whose digits sum to 9
Hard. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · Grok: 1 of 1 model passed, Grok 4.7 at $0.0056. The prompt in full.
GPT-6 Sol and Grok 4.7 both passed this one. GPT-6 Sol answered sooner (2.5 s to 11.1 s), and GPT-6 Sol's reply cost less ($0.0019 to $0.0056); GPT-6 Sol wrote 78 tokens of reply, Grok 4.7 145.
GPT-6 Sol
Passed: Final answer 165: right.
Subtract 1 from the thousands digit. The four digits then become nonnegative integers that sum to 8. By stars and bars, the number of solutions is \(\binom{8+4-1}{4-1}=\binom{11}{3}=165\). Every solution gives a valid four-digit integer. Final answer: 165378 tokens in, 117 out (39 of them reasoning) · 2.5 s · $0.0019 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Grok 4.7
Passed: Final answer 165: right.
The four-digit positive integers \(ABCD\) (with \(A \in \{1,\dots,9\}\) and \(B,C,D \in \{0,\dots,9\}\)) whose digits sum to 9 are in bijection with the non-negative integer solutions of \(A' + B + C + D = 8\) via the change of variables \(A' = A - 1\). No variable can exceed 8, so the upper bounds of 9 are automatically satisfied. The number of solutions is therefore \[ \binom{8+4-1}{3} = \binom{11}{3} = 165. \] Final answer: 1651,613 tokens in, 925 out (780 of them reasoning) · 11.1 s · $0.0056 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
5. The highest of three dice is a 5
Hard. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · Grok: 1 of 1 model passed, Grok 4.7 at $0.0050. The prompt in full.
GPT-6 Sol and Grok 4.7 both passed this one. GPT-6 Sol answered sooner (2.2 s to 9.1 s), and GPT-6 Sol's reply cost less ($0.0020 to $0.0050); GPT-6 Sol wrote 91 tokens of reply, Grok 4.7 165.
GPT-6 Sol
Passed: Final answer 61/216: right.
For the highest number to be exactly 5, all three dice must show 5 or less, but not all can show 4 or less. \[ P(\text{highest is }5)=\left(\frac56\right)^3-\left(\frac46\right)^3 =\frac{125-64}{216}=\frac{61}{216}. \] Final answer: 61/216393 tokens in, 122 out (31 of them reasoning) · 2.2 s · $0.0020 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Grok 4.7
Passed: Final answer 61/216: right.
The total number of outcomes when rolling three fair six-sided dice is \(6^3 = 216\). The highest face is exactly 5 precisely when every die shows a face in \(\{1,2,3,4,5\}\) and at least one die shows a 5. There are \(5^3 = 125\) outcomes in which every face is at most 5, and \(4^3 = 64\) outcomes in which every face is at most 4. Subtracting gives \[ 125 - 64 = 61 \] favorable outcomes. Thus the desired probability is \(\frac{61}{216}\). (The fraction is already in lowest terms, since 61 is prime and does not divide 216.) Final answer: 61/2161,628 tokens in, 786 out (621 of them reasoning) · 9.1 s · $0.0050 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
Each GPT model against each Grok model
Every GPT model against every Grok model on the same 5 math prompts: passes, the median wait and what the replies cost. A lead under 10% counts as close.
GPT-6 Astra vs Grok 4.7: 5 of 5 each; GPT-6 Astra answered 3.5× sooner at the median and Grok 4.7 cost 1.7× less. GPT-6 Astra vs Grok 4.7, on every job.
GPT-6 Sol vs Grok 4.7: 5 of 5 each; GPT-6 Sol answered 4.1× sooner at the median and cost 2.8× less. GPT-6 Sol vs Grok 4.7, on every job.
GPT-6 Luna vs Grok 4.7: GPT-6 Luna 4 of 5, Grok 4.7 5; GPT-6 Luna answered 4.1× sooner at the median and cost 51.1× less. GPT-6 Luna vs Grok 4.7, on every job.
How the math replies are scored
Every GPT and Grok reply above was checked the same way as every other model's, by the rules published with the prompts: how the math prompts are scored, and each one in full.
The differences at a glance
What follows from each model's facts in our catalog.
The lineups
GPT: 3 models, GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. Grok: one model, Grok 4.7.
Price per message
The least expensive GPT model is GPT-6 Luna (60 messages a day on Pro); Grok's one model is Grok 4.7 (250 messages a month on Pro).
Context window
GPT goes up to 1.05M tokens (GPT-6 Astra); Grok up to 500K tokens (Grok 4.7).
Images and PDFs
Every model here reads images. Every model here takes a PDF as a whole file.
On the Free plan
Free's one-time trial of 5 messages covers GPT-6 Sol, GPT-6 Luna, and Grok 4.7. Paid plans have every model, with messages every month.
Every GPT and Grok model's context window, files and API price: Grok vs ChatGPT.
Where your messages go
In llmwise, a message to GPT goes to its maker, OpenAI, or through OpenRouter when llmwise can't reach the maker directly. Grok models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.
Questions
Which is better, ChatGPT or Grok for math?
In our math test runs on September 27, 2026, GPT's 3 models passed 14 of 15; GPT-6 Astra and GPT-6 Sol each passed 5 of 5. Grok's one model passed 5 of 5 ($0.0050 a reply). GPT-6 Sol and Grok 4.7 each passed 5 of the 5 prompts, so these math prompts don't split GPT and Grok; the replies on this page show how they differ.
Is GPT or Grok cheaper?
In llmwise, the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro), and the least expensive Grok model is Grok 4.7 (250 messages a month on Pro). At API list prices (September 2026), a typical message of 4,000 tokens in and 700 out costs $0.0008 on GPT-6 Luna and $0.0098 on Grok 4.7.
Can I use GPT and Grok in the same chat?
Yes. Ask GPT-6 Sol a question, then switch the picker to Grok 4.7 and ask again: Grok 4.7 sees the whole conversation, GPT-6 Sol's answer included.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.