Comparison
Grok vs ChatGPT
In our test runs on October 2, 2026, the same 50 prompts across 10 jobs: Grok's one model passed 47 of 50 replies and GPT's 4 models passed 186 of 200 replies. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself.
Based on 250 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
Job by job, both families' best models shared the top result on 9 of the 10 jobs; GPT's alone had it on writing.
Grok vs GPT, job by job
On each job, Grok's pick against GPT's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.
Writing: Grok 4.7 passed 3 of 5 and GPT-6 Luna 5 of 5. GPT-6 Luna answered 2.5× sooner at the median, 6.0 s against 2.4 s. GPT-6 Luna cost 57.8× less, $0.0288 against $0.0005 for the 5 replies. Their replies ran to about the same length.
Coding: Grok 4.7 passed 5 of 5 and GPT-6 Luna 5 of 5. GPT-6 Luna answered 6.9× sooner at the median, 31.6 s against 4.6 s. GPT-6 Luna cost 85.8× less, $0.1229 against $0.0014 for the 5 replies. Grok 4.7's replies ran 21% longer, in tokens of reply, thinking not counted.
Math: Grok 4.7 passed 5 of 5 and GPT-6.1 Sol 5 of 5. GPT-6.1 Sol answered 6.6× sooner at the median, 11.7 s against 1.8 s. GPT-6.1 Sol cost 8.4× less, $0.0362 against $0.0043 for the 5 replies. Grok 4.7's replies ran 67% longer, in tokens of reply, thinking not counted.
Summarization: Grok 4.7 passed 5 of 5 and GPT-6.1 Sol 5 of 5. GPT-6.1 Sol answered 1.4× sooner at the median, 2.8 s against 2.0 s. GPT-6.1 Sol cost 2.8× less, $0.0141 against $0.0050 for the 5 replies. GPT-6.1 Sol's replies ran 18% longer, in tokens of reply, thinking not counted.
Data analysis: Grok 4.7 passed 5 of 5 and GPT-6 Luna 5 of 5. GPT-6 Luna answered 3.6× sooner at the median, 8.9 s against 2.5 s. GPT-6 Luna cost 63.2× less, $0.0556 against $0.0009 for the 5 replies. Grok 4.7's replies ran 65% longer, in tokens of reply, thinking not counted.
Customer support: Grok 4.7 passed 4 of 5 and GPT-6 Luna 4 of 5. GPT-6 Luna answered 8.3× sooner at the median, 11.6 s against 1.4 s. GPT-6 Luna cost 70.9× less, $0.0306 against $0.0004 for the 5 replies. Grok 4.7's replies ran 60% longer, in tokens of reply, thinking not counted.
Translation: Grok 4.7 passed 5 of 5 and GPT-6 Luna 5 of 5. GPT-6 Luna answered 6.9× sooner at the median, 10.7 s against 1.6 s. GPT-6 Luna cost 88.4× less, $0.0410 against $0.0005 for the 5 replies. Their replies ran to about the same length.
SQL: Grok 4.7 passed 5 of 5 and GPT-6 Luna 5 of 5. GPT-6 Luna answered 2.5× sooner at the median, 3.0 s against 1.2 s. GPT-6 Luna cost 54.0× less, $0.0240 against $0.0004 for the 5 replies. GPT-6 Luna's replies ran 10% longer, in tokens of reply, thinking not counted.
RAG and answering from documents: Grok 4.7 passed 5 of 5 and GPT-6 Luna 5 of 5. GPT-6 Luna answered 1.5× sooner at the median, 1.9 s against 1.3 s. GPT-6 Luna cost 37.7× less, $0.0146 against $0.0004 for the 5 replies. GPT-6 Luna's replies ran 25% longer, in tokens of reply, thinking not counted.
Agents and tool use: Grok 4.7 passed 5 of 5 and GPT-6 Luna 5 of 5. GPT-6 Luna answered 2.7× sooner at the median, 5.2 s against 1.9 s. GPT-6 Luna cost 48.5× less, $0.0197 against $0.0004 for the 5 replies. Their replies ran to about the same length.
The 2 prompts only one of Grok and GPT passed
Where one family's pick passed a prompt and the other's didn't, in each check's own words.
Rewrite corporate jargon in plain words (writing): GPT-6 Luna passed and Grok 4.7 didn't. Grok 4.7: Graded 3.7 of 5 on average (lowest 3). GPT-6 Luna: Graded 4.7 of 5 on average (lowest 4).
Argue both sides of free buses (writing): GPT-6 Luna passed and Grok 4.7 didn't. Grok 4.7: Graded 4.7 of 5 on average (lowest 4); but a paragraph of 100 words, over the 90 allowed. GPT-6 Luna: Graded 4.3 of 5 on average (lowest 4).
Each Grok model against each GPT model
Every Grok model against every GPT model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.
Grok 4.7 vs GPT-6 Astra: 47 and 48 of 50; 3 prompts split them; Grok 4.7's replies cost 1.5× less in all, and GPT-6 Astra answered sooner on 41, Grok 4.7 on 8. GPT-6 Astra vs Grok 4.7.
Grok 4.7 vs GPT-6.1 Sol: 47 and 46 of 50; 3 prompts split them; GPT-6.1 Sol's replies cost 6.5× less in all, and GPT-6.1 Sol answered sooner on 48, Grok 4.7 on 2.
Grok 4.7 vs GPT-6 Sol: 47 and 45 of 50; 4 prompts split them; GPT-6 Sol's replies cost 3.0× less in all, and GPT-6 Sol answered sooner on 49, Grok 4.7 on 0. GPT-6 Sol vs Grok 4.7.
Grok 4.7 vs GPT-6 Luna: 47 and 47 of 50; 4 prompts split them; GPT-6 Luna's replies cost 65.7× less in all, and GPT-6 Luna answered sooner on 49, Grok 4.7 on 0. GPT-6 Luna vs Grok 4.7.
One writing prompt, both replies
Grok 4.7 and GPT-6 Luna on a writing prompt, the job where they differed most: both replies as they came.
An everyday prompt: “Rewrite corporate jargon in plain words”, in full.
Grok 4.7
Failed: Graded 3.7 of 5 on average (lowest 3).
In the third quarter, teams from different parts of the company worked together to put a customer-focused approach into practice. Retention measures rose 12%, and the rate of customer loss fell by 30 basis points among enterprise customers. Contact points across all channels were also resized to use capacity more efficiently.
1,658 tokens in, 1,603 out (1,543 of them reasoning) · 20.0 s · $0.0112 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
GPT-6 Luna
Passed: Graded 4.7 of 5 on average (lowest 4).
In Q3, teams across the company worked together to focus more on customers. Retention measures improved by 12%, and the rate of customer loss fell by 0.3 percentage points in the enterprise segment. The company also adjusted its customer contact channels to make better use of available capacity.
417 tokens in, 124 out (59 of them reasoning) · 2.4 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·
Every model, every job
All 5 Grok and GPT models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.
| Model | Passed | Hard ones | Cost per reply | On Pro |
|---|---|---|---|---|
| Grok 4.7xAI | 47 of 50 | 19 of 20 | $0.0077 | Up to 250 a month |
| GPT-6 AstraOpenAI | 48 of 50 | 20 of 20 | $0.0117 | Up to 31 a month |
| GPT-6.1 SolOpenAI | 46 of 50 | 19 of 20 | $0.0012 | Up to 125 a month |
| GPT-6 SolOpenAI | 45 of 50 | 19 of 20 | $0.0026 | Up to 125 a month |
| GPT-6 LunaOpenAI | 47 of 50 | 20 of 20 | $0.00012 | Up to 60 a day |
Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
Their own subscriptions
Each company's own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.
SuperGrok (SpaceXAI), $30 a month
Models it names: Grok 4.6. On its limits: “Smarter answers in Expert mode: Access our best model with higher limits” SuperGrok vs llmwise.
Checked : Grok plans.
ChatGPT Plus (OpenAI), $20 a month
Models it names: GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. On its limits: “To ensure a smooth experience for all users, Plus subscriptions may include usage limits such as message caps, especially during high demand. These limits may vary based on system conditions.” ChatGPT Plus vs llmwise.
Checked : OpenAI Help Center: What is ChatGPT Plus?, OpenAI Help Center: Managing usage with GPT-6 Astra in Work and Codex and ChatGPT pricing.
llmwise Pro, $20 a month, has all 5 of these models in one chat, on one monthly allowance. On it: Grok 4.7 up to 250 messages a month and GPT-6.1 Sol up to 125.
The lineups at a glance
What follows from each model's facts in our catalog.
The lineups
Grok: one model, Grok 4.7. GPT: 4 models, GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna.
Price per message
Grok's one model is Grok 4.7 (250 messages a month on Pro); the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro).
Context window
Grok goes up to 500K tokens (Grok 4.7); GPT up to 1.05M tokens (GPT-6 Astra).
Images and PDFs
Every model here reads images. Every model here takes a PDF as a whole file.
On the Free plan
Free's one-time trial of 5 messages covers Grok 4.7, GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna. Paid plans have every model, with messages every month.
Model by model
Two named models side by side, prompt by prompt, each with its messages on every plan.
More head-to-heads
Each of Grok and GPT against the other families, every job from the same test runs.
Where your messages go
In llmwise, a message to GPT goes to its maker, OpenAI, or through OpenRouter when llmwise can't reach the maker directly. Grok models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.
Questions
Which is better, Grok or ChatGPT?
In our test runs on October 2, 2026, the same 50 prompts across 10 jobs: Grok's one model passed 47 of 50 replies and GPT's 4 models passed 186 of 200 replies. Job by job, both families' best models shared the top result on 9 of the 10 jobs; GPT's alone had it on writing.
Which is cheaper, Grok or GPT?
In llmwise, the least expensive Grok model is Grok 4.7 (250 messages a month on Pro), and the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro). At API list prices (October 2026), a typical message of 4,000 tokens in and 700 out costs $0.0122 on Grok 4.7 and $0.0008 on GPT-6 Luna.
Can I use Grok and GPT in the same chat?
Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.