Skip to content

Model vs model

Gemini 3.8 Flash vs Grok 4.7

Gemini 3.8 Flash and Grok 4.7 are both in llmwise. What each gets on every plan, what it reads, how it's served and what happens when its provider fails, from the catalog and the code that runs them.

Model prices and specs checked against OpenRouter's Gemini 3.8 Flash page, OpenRouter's Grok 4.7 page. Updated .

Short answer

They cost the same in llmwise: up to 250 messages a month on Pro on either. Beyond that, they read the same files and neither is easier to try. In our test runs, Gemini 3.8 Flash passed 46 of the 50 prompts both answered and Grok 4.7 47; 3 prompts split them, most on data analysis (4 to 5).

Gemini 3.8 Flash vs Grok 4.7, prompt by prompt

Every prompt Gemini 3.8 Flash and Grok 4.7 both answered, compared directly, their biggest differences first. One run each, through OpenRouter: a wait depends on the provider and the load that day, so a lead under 10% counts as close.

Of the 50 prompts both answered, both passed 45, only Gemini 3.8 Flash passed 1, only Grok 4.7 passed 2, and neither passed 2. Gemini 3.8 Flash answered sooner on 34 of the 50 and Grok 4.7 on 13; the rest were within 10% of each other. The 50 replies cost $0.0644 on Gemini 3.8 Flash and $0.3163 on Grok 4.7: 4.9× less on Gemini 3.8 Flash.

The 3 prompts only one of Gemini 3.8 Flash and Grok 4.7 passed

  • Announce a second bakery shop on LinkedIn (writing): Gemini 3.8 Flash passed and Grok 4.7 didn't. Gemini 3.8 Flash: Graded 4.5 of 5 on average (lowest 4). Grok 4.7: Graded 4.5 of 5 on average (lowest 4); but 94 words, under the 100 asked for.

  • A product announcement with five rules (writing): Grok 4.7 passed and Gemini 3.8 Flash didn't. Gemini 3.8 Flash: Graded 3.7 of 5 on average (lowest 3). Grok 4.7: Graded 4.0 of 5 on average (lowest 3).

  • Correlation between ad spend and sign-ups (data analysis): Grok 4.7 passed and Gemini 3.8 Flash didn't. Gemini 3.8 Flash: Final answer 0.91; expected 0.97. Grok 4.7: Final answer 0.97: right.

Job by job, the widest gaps first

  • Data analysis: Gemini 3.8 Flash passed 4 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 1.7× sooner at the median, 5.4 s against 9.4 s. Gemini 3.8 Flash cost 3.1× less, $0.0138 against $0.0423 for the 5 replies. Gemini 3.8 Flash's replies ran 265% longer, in tokens of reply, thinking not counted.

  • Coding: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 6.0× sooner at the median, 4.2 s against 25.3 s. Gemini 3.8 Flash cost 11.3× less, $0.0098 against $0.1099 for the 5 replies. Gemini 3.8 Flash's replies ran 44% longer, in tokens of reply, thinking not counted.

  • Translation: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 4.6× sooner at the median, 3.0 s against 13.8 s. Gemini 3.8 Flash cost 9.3× less, $0.0037 against $0.0347 for the 5 replies. Gemini 3.8 Flash's replies ran 19% longer, in tokens of reply, thinking not counted.

  • SQL: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 1.3× sooner at the median, 3.7 s against 4.9 s. Gemini 3.8 Flash cost 5.1× less, $0.0037 against $0.0189 for the 5 replies. Gemini 3.8 Flash's replies ran 26% longer, in tokens of reply, thinking not counted.

  • Summarization: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 1.3× sooner at the median, 3.7 s against 4.7 s. Gemini 3.8 Flash cost 4.3× less, $0.0038 against $0.0161 for the 5 replies. Their replies ran to about the same length.

  • Math: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Gemini 3.8 Flash answered 1.9× sooner at the median, 4.7 s against 9.1 s. Gemini 3.8 Flash cost 3.5× less, $0.0071 against $0.0251 for the 5 replies. Gemini 3.8 Flash's replies ran 41% longer, in tokens of reply, thinking not counted.

  • RAG and answering from documents: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Grok 4.7 answered 1.7× sooner at the median, 3.1 s against 1.9 s. Gemini 3.8 Flash cost 3.5× less, $0.0033 against $0.0115 for the 5 replies. Gemini 3.8 Flash's replies ran 73% longer, in tokens of reply, thinking not counted.

  • Writing: Gemini 3.8 Flash passed 3 of 5 and Grok 4.7 3 of 5. Their median waits were close, 3.9 s against 4.0 s. Gemini 3.8 Flash cost 3.2× less, $0.0066 against $0.0210 for the 5 replies. Gemini 3.8 Flash's replies ran 43% longer, in tokens of reply, thinking not counted.

  • Agents and tool use: Gemini 3.8 Flash passed 5 of 5 and Grok 4.7 5 of 5. Grok 4.7 answered 1.8× sooner at the median, 4.4 s against 2.5 s. Gemini 3.8 Flash cost 3.0× less, $0.0043 against $0.0129 for the 5 replies. Gemini 3.8 Flash's replies ran 10% longer, in tokens of reply, thinking not counted.

  • Customer support: Gemini 3.8 Flash passed 4 of 5 and Grok 4.7 4 of 5. Gemini 3.8 Flash answered 2.1× sooner at the median, 3.7 s against 7.9 s. Gemini 3.8 Flash cost 2.9× less, $0.0082 against $0.0239 for the 5 replies. Their replies ran to about the same length.

All 50 prompts: who passed, who answered sooner, who cost less
Gemini 3.8 Flash and Grok 4.7 on each prompt of our test runs
PromptResultSoonerCheaper
Turn a title into a URL slugBoth passedGemini 3.8 Flash, 1.6×took 4.2 s and 6.7 sGemini 3.8 Flash, 5.6×cost $0.0009 and $0.0048
Parse a duration like “1h 30m”Both passedGemini 3.8 Flash, 11.0×took 2.3 s and 25.3 sGemini 3.8 Flash, 14.6×cost $0.0010 and $0.0151
Merge overlapping intervalsBoth passedClosetook 3.8 s and 3.8 sGemini 3.8 Flash, 2.0×cost $0.0012 and $0.0025
Evaluate an arithmetic expression, no evalBoth passedGemini 3.8 Flash, 18.5×took 6.7 s and 124.9 sGemini 3.8 Flash, 10.6×cost $0.0045 and $0.0472
Parse CSV with quoted fieldsBoth passedGemini 3.8 Flash, 20.2×took 5.0 s and 100.4 sGemini 3.8 Flash, 18.4×cost $0.0022 and $0.0403
Announce a second bakery shop on LinkedInOnly Gemini 3.8 FlashClosetook 3.9 s and 4.0 sGemini 3.8 Flash, 2.4×cost $0.0009 and $0.0022
Rewrite corporate jargon in plain wordsNeither passedGemini 3.8 Flash, 6.1×took 3.1 s and 19.3 sGemini 3.8 Flash, 13.8×cost $0.0006 and $0.0088
Decline a meeting and offer two timesBoth passedGrok 4.7, 1.5×took 3.5 s and 2.3 sGemini 3.8 Flash, 2.4×cost $0.0008 and $0.0019
A product announcement with five rulesOnly Grok 4.7Grok 4.7, 2.9×took 7.6 s and 2.6 sGrok 4.7, 1.9×cost $0.0031 and $0.0017
Argue both sides of free busesBoth passedGemini 3.8 Flash, 2.3×took 5.2 s and 12.1 sGemini 3.8 Flash, 5.7×cost $0.0011 and $0.0065
A discount, then sales taxBoth passedGemini 3.8 Flash, 1.6×took 2.5 s and 3.9 sGemini 3.8 Flash, 4.1×cost $0.0007 and $0.0027
Pens at 3 for $4Both passedGemini 3.8 Flash, 8.7×took 4.7 s and 40.7 sGemini 3.8 Flash, 5.2×cost $0.0017 and $0.0087
Compound interest over three yearsBoth passedGemini 3.8 Flash, 2.6×took 2.1 s and 5.5 sGemini 3.8 Flash, 3.6×cost $0.0009 and $0.0031
Four-digit numbers whose digits sum to 9Both passedGemini 3.8 Flash, 2.4×took 4.7 s and 11.1 sGemini 3.8 Flash, 2.6×cost $0.0022 and $0.0056
The highest of three dice is a 5Both passedGemini 3.8 Flash, 1.9×took 4.9 s and 9.1 sGemini 3.8 Flash, 2.8×cost $0.0018 and $0.0050
An article in three bulletsBoth passedGemini 3.8 Flash, 2.7×took 3.7 s and 9.8 sGemini 3.8 Flash, 7.0×cost $0.0007 and $0.0051
An email thread in one sentenceBoth passedGrok 4.7, 2.2×took 3.8 s and 1.7 sGemini 3.8 Flash, 3.4×cost $0.0005 and $0.0018
Decisions and action items from a meetingBoth passedGemini 3.8 Flash, 1.5×took 3.2 s and 4.7 sGemini 3.8 Flash, 3.0×cost $0.0009 and $0.0028
A quarterly memo for the CEOBoth passedGemini 3.8 Flash, 1.5×took 4.0 s and 6.1 sGemini 3.8 Flash, 4.8×cost $0.0009 and $0.0044
A study with a negative resultBoth passedGrok 4.7, 1.2×took 3.6 s and 3.0 sGemini 3.8 Flash, 3.1×cost $0.0007 and $0.0020
The region with the most revenueBoth passedGemini 3.8 Flash, 1.3×took 7.4 s and 9.4 sGemini 3.8 Flash, 1.5×cost $0.0038 and $0.0057
Average order value in AugustBoth passedGemini 3.8 Flash, 1.3×took 4.4 s and 5.5 sGemini 3.8 Flash, 2.3×cost $0.0016 and $0.0037
Revenue change from July to AugustBoth passedGemini 3.8 Flash, 1.5×took 6.6 s and 10.1 sGemini 3.8 Flash, 1.9×cost $0.0037 and $0.0071
A median, filtered two waysBoth passedGemini 3.8 Flash, 1.2×took 3.5 s and 4.2 sGemini 3.8 Flash, 3.4×cost $0.0010 and $0.0036
Correlation between ad spend and sign-upsOnly Grok 4.7Gemini 3.8 Flash, 8.4×took 5.4 s and 45.4 sGemini 3.8 Flash, 6.0×cost $0.0037 and $0.0222
A late orderBoth passedGemini 3.8 Flash, 5.5×took 3.1 s and 17.2 sGemini 3.8 Flash, 8.2×cost $0.0009 and $0.0076
A return inside the windowBoth passedGemini 3.8 Flash, 1.6×took 2.9 s and 4.7 sGemini 3.8 Flash, 3.7×cost $0.0009 and $0.0032
A frustrated customerNeither passedGemini 3.8 Flash, 2.6×took 3.7 s and 9.4 sGemini 3.8 Flash, 4.7×cost $0.0011 and $0.0050
A refund request outside the windowBoth passedGemini 3.8 Flash, 1.3×took 5.9 s and 7.6 sGemini 3.8 Flash, 1.4×cost $0.0029 and $0.0041
A message with a planted instructionBoth passedClosetook 7.3 s and 7.9 sGemini 3.8 Flash, 1.6×cost $0.0025 and $0.0040
A delivery message into SpanishBoth passedGemini 3.8 Flash, 3.7×took 3.7 s and 13.8 sGemini 3.8 Flash, 8.4×cost $0.0008 and $0.0066
A product description into FrenchBoth passedGemini 3.8 Flash, 10.4×took 2.2 s and 22.7 sGemini 3.8 Flash, 12.9×cost $0.0008 and $0.0097
A meeting note into GermanBoth passedGemini 3.8 Flash, 4.2×took 3.0 s and 12.6 sGemini 3.8 Flash, 9.5×cost $0.0006 and $0.0060
Idioms into natural JapaneseBoth passedGemini 3.8 Flash, 2.1×took 4.7 s and 10.0 sGemini 3.8 Flash, 6.2×cost $0.0008 and $0.0048
A lease clause into Brazilian PortugueseBoth passedGemini 3.8 Flash, 7.3×took 2.0 s and 15.0 sGemini 3.8 Flash, 9.6×cost $0.0008 and $0.0076
Customers in one countryBoth passedGrok 4.7, 1.7×took 3.1 s and 1.8 sGemini 3.8 Flash, 3.8×cost $0.0005 and $0.0018
Count orders by statusBoth passedGrok 4.7, 2.6×took 3.7 s and 1.4 sGemini 3.8 Flash, 3.6×cost $0.0005 and $0.0017
Revenue by categoryBoth passedGemini 3.8 Flash, 1.6×took 3.0 s and 4.9 sGemini 3.8 Flash, 4.1×cost $0.0008 and $0.0031
Every customer, even those without ordersBoth passedGemini 3.8 Flash, 1.5×took 3.7 s and 5.3 sGemini 3.8 Flash, 4.7×cost $0.0008 and $0.0037
Monthly revenue with a running totalBoth passedGemini 3.8 Flash, 3.6×took 4.7 s and 16.9 sGemini 3.8 Flash, 7.0×cost $0.0012 and $0.0086
A fact from one sectionBoth passedGrok 4.7, 1.2×took 3.1 s and 2.7 sGemini 3.8 Flash, 4.2×cost $0.0006 and $0.0025
Core hours and start timesBoth passedGrok 4.7, 1.6×took 2.9 s and 1.9 sGemini 3.8 Flash, 3.4×cost $0.0006 and $0.0022
Two sections in one answerBoth passedGrok 4.7, 2.2×took 3.1 s and 1.5 sGemini 3.8 Flash, 2.7×cost $0.0006 and $0.0018
A later amendment changes the answerBoth passedGemini 3.8 Flash, 1.4×took 4.1 s and 5.5 sGemini 3.8 Flash, 3.4×cost $0.0009 and $0.0030
A question the handbook doesn't answerBoth passedGrok 4.7, 1.6×took 2.9 s and 1.8 sGemini 3.8 Flash, 3.8×cost $0.0006 and $0.0021
Pick the tool and work out the dateBoth passedGrok 4.7, 1.2×took 1.9 s and 1.5 sGemini 3.8 Flash, 3.6×cost $0.0005 and $0.0018
Convert a currencyBoth passedGrok 4.7, 2.8×took 4.4 s and 1.5 sGemini 3.8 Flash, 1.8×cost $0.0008 and $0.0013
Book a meeting from a sentenceBoth passedGemini 3.8 Flash, 1.4×took 6.3 s and 9.1 sGemini 3.8 Flash, 4.0×cost $0.0012 and $0.0048
Search, but don't bookBoth passedGrok 4.7, 2.6×took 6.3 s and 2.5 sGemini 3.8 Flash, 2.2×cost $0.0009 and $0.0020
Two calls with a unit conversionBoth passedGemini 3.8 Flash, 1.2×took 3.9 s and 4.5 sGemini 3.8 Flash, 3.0×cost $0.0010 and $0.0029

Gemini 3.8 Flash vs Grok 4.7 in our test runs

Gemini 3.8 Flash and Grok 4.7 on the same prompts, job by job: how many replies passed their check.

Based on 100 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Gemini 3.8 Flash and Grok 4.7 in our test runs, job by job
JobGemini 3.8 FlashGrok 4.7
Coding5 of 55 of 5
Writing3 of 53 of 5
Math5 of 55 of 5
Summarization5 of 55 of 5
Data analysis4 of 55 of 5
Customer support4 of 54 of 5
Translation5 of 55 of 5
SQL5 of 55 of 5
RAG and answering from documents5 of 55 of 5
Agents and tool use5 of 55 of 5
All jobs46 of 50 · 1 message each · $0.0013 a reply47 of 50 · 1 message each · $0.0063 a reply

The same prompts, sent to both through llmwise's own pipeline and checked the same way: five for each job. A bold count is the one that passed more. Messages: what each reply counted as on Pro. Cost: what OpenRouter charged us per reply, on average. Every prompt, and how it's scored.

Where they differed most, data analysis: both replies to a prompt only one of them passed.

A hard prompt: “Correlation between ad spend and sign-ups”, in full.

  • Gemini 3.8 Flash

    Failed: Final answer 0.91; expected 0.97.

    Let $x$ represent `ad_spend` and $y$ represent `signups`. The number of data points is $n = 8$.
    
    **Data:**
    - $x = [500, 800, 650, 1200, 900, 300, 1100, 700]$
    - $y = [42, 55, 49, 71, 60, 35, 64, 58]$
    
    **Summary Calculations:**
    - $\sum x = 500 + 800 + 650 + 1200 + 900 + 300 + 1100 + 700 = 6150$
    - $\sum y = 42 + 55 + 49 + 71 + 60 + 35 + 64 + 58 = 434$
    - $\sum x^2 = 500^2 + 800^2 + 650^2 + 1200^2 + 900^2 + 300^2 + 1100^2 + 700^2 = 5,342,500$
    - $\sum y^2 = 42^2 + 55^2 + 49^2 + 71^2 + 60^2 + 35^2 + 64^2 + 58^2 = 24,676$
    - $\sum xy = (500)(42) + (800)(55) + (650)(49) + (1200)(71) + (900)(60) + (300)(35) + (1100)(64) + (700)(58)$
      $= 21,000 + 44,000 + 31,850 + 85,200 + 54,000 + 10,500 + 70,400 + 40,600 = 357,550$
    
    **Pearson Correlation Formula:**
    $$r = \frac{n\sum xy - (\sum x)(\sum y)}{\sqrt{[n\sum x^2 - (\sum x)^2][n\sum y^2 - (\sum y)^2]}}$$
    
    - **Numerator:**
      $8(357,550) - (6150)(434) = 2,860,400 - 2,669,100 = 191,300$
    
    - **Denominator:**
      $n\sum x^2 - (\sum x)^2 = 8(5,342,500) - 6150^2 = 42,740,000 - 37,822,500 = 4,917,500$
      $n\sum y^2 - (\sum y)^2 = 8(24,676) - 434^2 = 197,408 - 188,356 = 9,052$
      $\sqrt{4,917,500 \times 9,052} = \sqrt{44,513,210,000} \approx 210,981.54$
    
    - **Correlation ($r$):**
      $r = \frac{191,300}{210,981.54} \approx 0.9067 \approx 0.91$
    
    Final answer: 0.91

    505 tokens in, 887 out · 5.4 s · $0.0037 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

  • Grok 4.7

    Passed: Final answer 0.97: right.

    **Pearson r (computational formula)**
    
    n = 8, x = ad_spend, y = signups
    
    - Σx = 6150, Σy = 434, Σxy = 357550, Σx² = 5352500, Σy² = 24516
    
    Numerator: nΣxy − ΣxΣy = 8×357550 − 6150×434 = 191300
    
    Denominator parts: nΣx² − (Σx)² = 4997500, nΣy² − (Σy)² = 7772
    
    r = 191300 / √(4997500 × 7772) = 191300 / √38840570000 ≈ 191300 / 197080.11 ≈ 0.9707
    
    **Final answer: 0.97**

    1,686 tokens in, 4,422 out (4,245 of them reasoning) · 45.4 s · $0.0222 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

Gemini 3.8 Flash and Grok 4.7 on every plan

Whether the one-time free trial reaches each model, then each paid plan's messages on it.

Messages on Gemini 3.8 Flash and Grok 4.7, plan by plan
PlanPriceGemini 3.8 FlashGrok 4.7
Free$0In the one-time trial of 5 messagesIn the one-time trial of 5 messages
Pro$20 a monthUp to 250 a monthUp to 250 a month
Max$50 a monthUp to 800 a monthUp to 800 a month
Ultra$100 a monthUp to 1,800 a monthUp to 1,800 a month
Studio$200 a monthUp to 4,000 a monthUp to 4,000 a month

Prices don't include tax, which is added where it applies and shown before you pay. A paid plan's month is one allowance shared by every model, so each monthly count is the most you get if all of it goes to that model. It renews each billing period; everyday models refill daily at 00:00 UTC. Long chats count more per reply. How pricing works.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

What differs

  • Messages on Pro

    They cost the same in llmwise: up to 250 messages a month on Pro on either.

  • Context window

    Gemini 3.8 Flash takes up to 1.05M tokens; Grok 4.7 up to 500K tokens. A chat in llmwise holds up to 200k tokens, which fits in either, so the difference shows only through each maker's own API.

  • Images and PDFs

    Both read images. Both take a PDF as the whole file, pages and all.

  • On Free

    Both are in the free trial.

  • Where messages go

    Gemini 3.8 Flash: Sent to Google directly. Grok 4.7: Served through OpenRouter by xAI alone, on an endpoint that doesn't store or train on prompts.

Fact by fact

Gemini 3.8 Flash and Grok 4.7, fact by fact
FactGemini 3.8 FlashGrok 4.7
Context window1.05M tokens500K tokens
Reads imagesYesYes
PDFsWhole fileWhole file
ReasoningYesYes
API price (September 2026)$0.75 in / $3.75 out per million tokens$1.60 in / $4.80 out per million tokens
A typical message at API prices (4,000 tokens in, 700 out)$0.0056$0.0098
A $10 top-up adds200 messages200 messages
Where a message goesSent to Google directly.Served through OpenRouter by xAI alone, on an endpoint that doesn't store or train on prompts.
If the provider failsIf Google fails before the reply starts (an overload, a server error, a dropped connection), llmwise sends the same request to Gemini 3.8 Flash through OpenRouter instead.xAI is its only host, so there's no other host to move to: if xAI fails, send the message again or pick another model.
Anthropic's safety fallbackDoesn't applyDoesn't apply
API prices are what our model catalog lists (Gemini: Google's list price; Grok: xAI's price, served through OpenRouter). In llmwise you pay per message, not per token: the counts above are what you get.

Gemini 3.8 Flash or Grok 4.7?

From the facts above and our test runs: the rest is how their answers suit your work, which one chat can show you.

  • Pick Gemini 3.8 Flash: it costs its maker less to run ($0.0056 a typical message at API prices), though in llmwise the count is the same.

Each model's page, the families, and other pairs

Gemini 3.8 Flash vs Grok 4.7 is one pair of models. The page below covers the whole families.

Questions

Is Gemini 3.8 Flash or Grok 4.7 cheaper in llmwise?

They cost the same in llmwise: up to 250 messages a month on Pro on either. Every paid plan's monthly allowance is shared by all models, so each count is the most you get if it all goes to that model.

Can I try Gemini 3.8 Flash and Grok 4.7 for free?

Yes: both are in the free trial of 5 messages.

Which has the bigger context window, Gemini 3.8 Flash or Grok 4.7?

Gemini 3.8 Flash: 1.05M tokens, against 500K tokens. A chat in llmwise holds up to 200k tokens, which fits in either, so the difference shows only through each maker's own API.

Can I use Gemini 3.8 Flash and Grok 4.7 in the same chat?

Yes. Pick Gemini 3.8 Flash for one message and Grok 4.7 for the next; the second sees the whole chat, including the first one's answer.

Which did better in your test runs, Gemini 3.8 Flash or Grok 4.7?

On the same 50 prompts, run on September 27, 2026, Gemini 3.8 Flash passed 46 and Grok 4.7 passed 47. The table on this page has each job, and every reply is published.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.