Comparison
Gemini vs Grok
On llmwise Pro, Gemini 3.1 Pro (preview) up to 125 messages a month and Grok 4.7 up to 250. We ran the same 50 prompts across 10 jobs on every Gemini and Grok model and published every reply: each family's pick against the other's job by job, the prompts where they split, then their plans and how the lineups differ.
Based on 150 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: Gemini's 2 models passed 95 of 100 replies and Grok's one model passed 47 of 50 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; Gemini's alone had it on writing and customer support.
Gemini vs Grok, job by job
On each job, Gemini's pick against Grok's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.
Writing: Gemini 3.1 Pro (preview) passed 4 of 5 and Grok 4.7 3 of 5. Grok 4.7 answered 2.3× sooner at the median, 9.1 s against 4.0 s. Grok 4.7 cost 2.2× less, $0.0459 against $0.0210 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 38% longer, in tokens of reply, thinking not counted.
Customer support: Gemini 3.1 Pro (preview) passed 5 of 5 and Grok 4.7 4 of 5. Grok 4.7 answered 1.1× sooner at the median, 8.7 s against 7.9 s. Grok 4.7 cost 2.2× less, $0.0529 against $0.0239 for the 5 replies. Their replies ran to about the same length.
Coding: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 6.0× sooner at the median, 4.2 s against 25.3 s. Gemini 3.8 Flash cost 11.3× less, $0.0098 against $0.1099 for the 5 replies. Gemini 3.8 Flash's replies ran 44% longer, in tokens of reply, thinking not counted.
Math: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 1.9× sooner at the median, 4.7 s against 9.1 s. Gemini 3.8 Flash cost 3.5× less, $0.0071 against $0.0251 for the 5 replies. Gemini 3.8 Flash's replies ran 41% longer, in tokens of reply, thinking not counted.
Summarization: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 1.3× sooner at the median, 3.7 s against 4.7 s. Gemini 3.8 Flash cost 4.3× less, $0.0038 against $0.0161 for the 5 replies. Their replies ran to about the same length.
Data analysis: Gemini 3.1 Pro (preview) passed 5 of 5 and Grok 4.7 5 of 5. Their median waits were close, 8.6 s against 9.4 s. Grok 4.7 cost 1.8× less, $0.0773 against $0.0423 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 178% longer, in tokens of reply, thinking not counted.
Translation: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 4.6× sooner at the median, 3.0 s against 13.8 s. Gemini 3.8 Flash cost 9.3× less, $0.0037 against $0.0347 for the 5 replies. Gemini 3.8 Flash's replies ran 19% longer, in tokens of reply, thinking not counted.
SQL: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 1.3× sooner at the median, 3.7 s against 4.9 s. Gemini 3.8 Flash cost 5.1× less, $0.0037 against $0.0189 for the 5 replies. Gemini 3.8 Flash's replies ran 26% longer, in tokens of reply, thinking not counted.
RAG and answering from documents: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Grok 4.7 answered 1.7× sooner at the median, 3.1 s against 1.9 s. Gemini 3.8 Flash cost 3.5× less, $0.0033 against $0.0115 for the 5 replies. Gemini 3.8 Flash's replies ran 73% longer, in tokens of reply, thinking not counted.
Agents and tool use: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Grok 4.7 answered 1.8× sooner at the median, 4.4 s against 2.5 s. Gemini 3.8 Flash cost 3.0× less, $0.0043 against $0.0129 for the 5 replies. Gemini 3.8 Flash's replies ran 10% longer, in tokens of reply, thinking not counted.
The 4 prompts only one of Gemini and Grok passed
Where one family's pick passed a prompt and the other's didn't, in each check's own words.
Announce a second bakery shop on LinkedIn (writing): Gemini 3.1 Pro (preview) passed and Grok 4.7 didn't. Gemini 3.1 Pro (preview): Graded 4.0 of 5 on average (lowest 3). Grok 4.7: Graded 4.5 of 5 on average (lowest 4); but 94 words, under the 100 asked for.
Rewrite corporate jargon in plain words (writing): Gemini 3.1 Pro (preview) passed and Grok 4.7 didn't. Gemini 3.1 Pro (preview): Graded 4.0 of 5 on average (lowest 3). Grok 4.7: Graded 3.7 of 5 on average (lowest 3).
Argue both sides of free buses (writing): Grok 4.7 passed and Gemini 3.1 Pro (preview) didn't. Gemini 3.1 Pro (preview): Graded 3.7 of 5 on average (lowest 3). Grok 4.7: Graded 4.3 of 5 on average (lowest 4).
A frustrated customer (customer support): Gemini 3.1 Pro (preview) passed and Grok 4.7 didn't. Gemini 3.1 Pro (preview): Graded 4.0 of 5 on average (lowest 3). Grok 4.7: Graded 3.3 of 5 on average (lowest 3).
Each Gemini model against each Grok model
Every Gemini model against every Grok model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.
Gemini 3.1 Pro vs Grok 4.7: 49 and 47 of 50; 4 prompts split them; Grok 4.7's replies cost 1.7× less in all, and Grok 4.7 answered sooner on 26, Gemini 3.1 Pro on 18. Gemini 3.1 Pro (preview) vs Grok 4.7.
Gemini 3.8 Flash vs Grok 4.7: 46 and 47 of 50; 3 prompts split them; Gemini 3.8 Flash's replies cost 4.9× less in all, and Gemini 3.8 Flash answered sooner on 34, Grok 4.7 on 13. Gemini 3.8 Flash vs Grok 4.7.
One writing prompt, both replies
Gemini 3.1 Pro (preview) and Grok 4.7 on a writing prompt, the job where they differed most: both replies as they came.
An everyday prompt: “Announce a second bakery shop on LinkedIn”, in full.
Gemini 3.1 Pro
Passed: Graded 4.0 of 5 on average (lowest 3).
We are absolutely thrilled to share some wonderful news with you all. Our little bakery, Crumb & Co., is growing! Thanks to your incredible support over the years, we are opening the doors to our second shop in Leeds on the 14th of November. When we first started baking, we never imagined we would be able to share our pastries and breads with so many lovely people. This new chapter in Leeds means the world to us, and we cannot wait to become a part of the local community. To celebrate this exciting milestone, we would love for you to join us on our opening day. Pop in to say hello, check out the new space, and enjoy a free cup of coffee on us. We are so grateful for your continued support and look forward to welcoming you to our new home very soon.
417 tokens in, 876 out (707 of them reasoning) · 9.5 s · $0.0113 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·
Grok 4.7
Failed: Graded 4.5 of 5 on average (lowest 4); but 94 words, under the 100 asked for.
We're opening a second Crumb & Co. shop in Leeds on 14 November. The new bakery will offer the same bread, cakes and pastries you already know from our first shop, baked fresh each morning. If you've been asking for a Leeds location, this one is for you. On opening day, coffee is free with any purchase. Come in, say hello, and try something from the counter. We'd love to meet you and hear what you'd like to see on the shelves. Thank you for supporting a small bakery. We can't wait to welcome you.
1,624 tokens in, 207 out (93 of them reasoning) · 4.0 s · $0.0022 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
Every model, every job
All 3 Gemini and Grok models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.
| Model | Passed | Hard ones | Cost per reply | On Pro |
|---|---|---|---|---|
| Gemini 3.1 Pro (preview)Google | 49 of 50 | 19 of 20 | $0.0107 | Up to 125 a month |
| Gemini 3.8 FlashGoogle | 46 of 50 | 18 of 20 | $0.0013 | Up to 250 a month |
| Grok 4.7xAI | 47 of 50 | 20 of 20 | $0.0063 | Up to 250 a month |
Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
Their own subscriptions
Each company's own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.
Google AI Pro (Google), $19.99 a month
Models it names: Gemini 3.1 Pro. On its limits: “AI Pro: 4x higher than standard limits” Google AI Pro vs llmwise.
Sources: Google One: Google AI plans and Gemini Apps Help: limits and upgrades for Google AI subscribers, checked .
SuperGrok (xAI), $30 a month
Models it names: Grok 4.6. On its limits: “Smarter answers in Expert mode: Access our best model with higher limits” SuperGrok vs llmwise.
Source: Grok plans, checked .
llmwise Pro, $20 a month, has all 3 of these models in one chat, from one monthly allowance: Gemini 3.1 Pro (preview) up to 125 messages a month and Grok 4.7 up to 250.
The lineups at a glance
What follows from each model's facts in our catalog.
The lineups
Gemini: 2 models, Gemini 3.1 Pro (preview) and Gemini 3.8 Flash. Grok: one model, Grok 4.7.
Price per message
The least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro); Grok's one model is Grok 4.7 (250 messages a month on Pro).
Context window
Gemini goes up to 1.05M tokens (Gemini 3.1 Pro (preview)); Grok up to 500K tokens (Grok 4.7).
Images and PDFs
Every model here reads images. Every model here takes a PDF as a whole file.
On the Free plan
Free's one-time trial of 5 messages covers Gemini 3.1 Pro (preview), Gemini 3.8 Flash, and Grok 4.7. Paid plans have every model, with messages every month.
Model by model
Two named models side by side, prompt by prompt, each with its messages on every plan.
More head-to-heads
Each of Gemini and Grok against the other families, every job from the same test runs.
Where your messages go
In llmwise, a message to Gemini goes to its maker, Google, or through OpenRouter when llmwise can't reach the maker directly. Grok models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.
Questions
Which is better, Gemini or Grok?
In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: Gemini's 2 models passed 95 of 100 replies and Grok's one model passed 47 of 50 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; Gemini's alone had it on writing and customer support.
Which is cheaper, Gemini or Grok?
In llmwise, the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro), and the least expensive Grok model is Grok 4.7 (250 messages a month on Pro). At API list prices (September 2026), a typical message of 4,000 tokens in and 700 out costs $0.0056 on Gemini 3.8 Flash and $0.0098 on Grok 4.7.
Can I use Gemini and Grok in the same chat?
Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.