Skip to content

Model vs model

DeepSeek V4.1 Flash vs Grok 4.7

DeepSeek V4.1 Flash and Grok 4.7 are both in llmwise. What each gets on every plan, what it reads, how it's served and what happens when its provider fails, from the catalog and the code that runs them.

Model prices and specs checked against OpenRouter's DeepSeek V4.1 Flash page, OpenRouter's Grok 4.7 page. Updated .

Short answer

DeepSeek V4.1 Flash is an everyday model (60 messages a day on Pro); Grok 4.7 draws on the monthly allowance (up to 250 messages a month on Pro). Otherwise, only Grok 4.7 reads a PDF as the whole file. In our test runs, DeepSeek V4.1 Flash passed 49 of the 50 prompts both answered and Grok 4.7 47; 2 prompts split them, most on writing (4 to 3).

DeepSeek V4.1 Flash vs Grok 4.7, prompt by prompt

Every prompt DeepSeek V4.1 Flash and Grok 4.7 both answered, compared directly, their biggest differences first. One run each, through OpenRouter: a wait depends on the provider and the load that day, so a lead under 10% counts as close.

Of the 50 prompts both answered, both passed 47, only DeepSeek V4.1 Flash passed 2, only Grok 4.7 passed 0, and neither passed 1. DeepSeek V4.1 Flash answered sooner on 45 of the 50 and Grok 4.7 on 3; the rest were within 10% of each other. The 50 replies cost $0.0172 on DeepSeek V4.1 Flash and $0.3163 on Grok 4.7: 18.4× less on DeepSeek V4.1 Flash.

The 2 prompts only one of DeepSeek V4.1 Flash and Grok 4.7 passed

  • Announce a second bakery shop on LinkedIn (writing): DeepSeek V4.1 Flash passed and Grok 4.7 didn't. DeepSeek V4.1 Flash: Graded 5.0 of 5 on average (lowest 5). Grok 4.7: Graded 4.5 of 5 on average (lowest 4); but 94 words, under the 100 asked for.

  • A frustrated customer (customer support): DeepSeek V4.1 Flash passed and Grok 4.7 didn't. DeepSeek V4.1 Flash: Graded 4.7 of 5 on average (lowest 4). Grok 4.7: Graded 3.3 of 5 on average (lowest 3).

Job by job, the widest gaps first

  • Customer support: DeepSeek V4.1 Flash passed 5 of 5 and Grok 4.7 4 of 5. DeepSeek V4.1 Flash answered 3.1× sooner at the median, 2.5 s against 7.9 s. DeepSeek V4.1 Flash cost 16.6× less, $0.0014 against $0.0239 for the 5 replies. DeepSeek V4.1 Flash's replies ran 34% longer, in tokens of reply, thinking not counted.

  • Writing: DeepSeek V4.1 Flash passed 4 of 5 and Grok 4.7 3 of 5. DeepSeek V4.1 Flash answered 1.3× sooner at the median, 3.1 s against 4.0 s. DeepSeek V4.1 Flash cost 14.7× less, $0.0014 against $0.0210 for the 5 replies. DeepSeek V4.1 Flash's replies ran 31% longer, in tokens of reply, thinking not counted.

  • SQL: DeepSeek V4.1 Flash passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4.1 Flash answered 3.0× sooner at the median, 1.7 s against 4.9 s. DeepSeek V4.1 Flash cost 23.3× less, $0.0008 against $0.0189 for the 5 replies. DeepSeek V4.1 Flash's replies ran 11% longer, in tokens of reply, thinking not counted.

  • Translation: DeepSeek V4.1 Flash passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4.1 Flash answered 7.6× sooner at the median, 1.8 s against 13.8 s. DeepSeek V4.1 Flash cost 23.2× less, $0.0015 against $0.0347 for the 5 replies. DeepSeek V4.1 Flash's replies ran 81% longer, in tokens of reply, thinking not counted.

  • Math: DeepSeek V4.1 Flash passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4.1 Flash answered 5.4× sooner at the median, 1.7 s against 9.1 s. DeepSeek V4.1 Flash cost 21.7× less, $0.0012 against $0.0251 for the 5 replies. Their replies ran to about the same length.

  • Coding: DeepSeek V4.1 Flash passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4.1 Flash answered 3.3× sooner at the median, 7.8 s against 25.3 s. DeepSeek V4.1 Flash cost 20.3× less, $0.0054 against $0.1099 for the 5 replies. DeepSeek V4.1 Flash's replies ran 24% longer, in tokens of reply, thinking not counted.

  • Summarization: DeepSeek V4.1 Flash passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4.1 Flash answered 3.1× sooner at the median, 1.5 s against 4.7 s. DeepSeek V4.1 Flash cost 19.3× less, $0.0008 against $0.0161 for the 5 replies. Their replies ran to about the same length.

  • Data analysis: DeepSeek V4.1 Flash passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4.1 Flash answered 3.2× sooner at the median, 3.0 s against 9.4 s. DeepSeek V4.1 Flash cost 16.5× less, $0.0026 against $0.0423 for the 5 replies. DeepSeek V4.1 Flash's replies ran 55% longer, in tokens of reply, thinking not counted.

  • Agents and tool use: DeepSeek V4.1 Flash passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4.1 Flash answered 1.8× sooner at the median, 1.4 s against 2.5 s. DeepSeek V4.1 Flash cost 14.1× less, $0.0009 against $0.0129 for the 5 replies. DeepSeek V4.1 Flash's replies ran 10% longer, in tokens of reply, thinking not counted.

  • RAG and answering from documents: DeepSeek V4.1 Flash passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4.1 Flash answered 1.7× sooner at the median, 1.1 s against 1.9 s. DeepSeek V4.1 Flash cost 10.1× less, $0.0011 against $0.0115 for the 5 replies. DeepSeek V4.1 Flash's replies ran 160% longer, in tokens of reply, thinking not counted.

All 50 prompts: who passed, who answered sooner, who cost less
DeepSeek V4.1 Flash and Grok 4.7 on each prompt of our test runs
PromptResultSoonerCheaper
Turn a title into a URL slugBoth passedDeepSeek V4.1 Flash, 3.6×took 1.9 s and 6.7 sDeepSeek V4.1 Flash, 19.4×cost $0.0002 and $0.0048
Parse a duration like “1h 30m”Both passedDeepSeek V4.1 Flash, 3.3×took 7.8 s and 25.3 sDeepSeek V4.1 Flash, 18.2×cost $0.0008 and $0.0151
Merge overlapping intervalsBoth passedDeepSeek V4.1 Flash, 2.6×took 1.5 s and 3.8 sDeepSeek V4.1 Flash, 17.2×cost $0.0001 and $0.0025
Evaluate an arithmetic expression, no evalBoth passedDeepSeek V4.1 Flash, 12.0×took 10.4 s and 124.9 sDeepSeek V4.1 Flash, 16.6×cost $0.0028 and $0.0472
Parse CSV with quoted fieldsBoth passedDeepSeek V4.1 Flash, 8.6×took 11.6 s and 100.4 sDeepSeek V4.1 Flash, 29.9×cost $0.0013 and $0.0403
Announce a second bakery shop on LinkedInOnly DeepSeek V4.1 FlashDeepSeek V4.1 Flash, 1.4×took 2.9 s and 4.0 sDeepSeek V4.1 Flash, 9.9×cost $0.0002 and $0.0022
Rewrite corporate jargon in plain wordsNeither passedDeepSeek V4.1 Flash, 16.4×took 1.2 s and 19.3 sDeepSeek V4.1 Flash, 30.1×cost $0.0003 and $0.0088
Decline a meeting and offer two timesBoth passedGrok 4.7, 1.4×took 3.1 s and 2.3 sDeepSeek V4.1 Flash, 7.8×cost $0.0002 and $0.0019
A product announcement with five rulesBoth passedGrok 4.7, 2.4×took 6.2 s and 2.6 sDeepSeek V4.1 Flash, 4.0×cost $0.0004 and $0.0017
Argue both sides of free busesBoth passedDeepSeek V4.1 Flash, 3.1×took 3.9 s and 12.1 sDeepSeek V4.1 Flash, 25.3×cost $0.0003 and $0.0065
A discount, then sales taxBoth passedDeepSeek V4.1 Flash, 6.9×took 0.6 s and 3.9 sDeepSeek V4.1 Flash, 13.0×cost $0.0002 and $0.0027
Pens at 3 for $4Both passedDeepSeek V4.1 Flash, 1.6×took 25.6 s and 40.7 sDeepSeek V4.1 Flash, 28.3×cost $0.0003 and $0.0087
Compound interest over three yearsBoth passedDeepSeek V4.1 Flash, 3.8×took 1.5 s and 5.5 sDeepSeek V4.1 Flash, 15.0×cost $0.0002 and $0.0031
Four-digit numbers whose digits sum to 9Both passedDeepSeek V4.1 Flash, 6.6×took 1.7 s and 11.1 sDeepSeek V4.1 Flash, 22.6×cost $0.0002 and $0.0056
The highest of three dice is a 5Both passedDeepSeek V4.1 Flash, 3.5×took 2.5 s and 9.1 sDeepSeek V4.1 Flash, 26.7×cost $0.0002 and $0.0050
An article in three bulletsBoth passedDeepSeek V4.1 Flash, 3.9×took 2.5 s and 9.8 sDeepSeek V4.1 Flash, 17.0×cost $0.0003 and $0.0051
An email thread in one sentenceBoth passedDeepSeek V4.1 Flash, 1.2×took 1.4 s and 1.7 sDeepSeek V4.1 Flash, 15.8×cost $0.0001 and $0.0018
Decisions and action items from a meetingBoth passedDeepSeek V4.1 Flash, 3.1×took 1.5 s and 4.7 sDeepSeek V4.1 Flash, 14.8×cost $0.0002 and $0.0028
A quarterly memo for the CEOBoth passedDeepSeek V4.1 Flash, 1.9×took 3.2 s and 6.1 sDeepSeek V4.1 Flash, 62.6×cost $0.0001 and $0.0044
A study with a negative resultBoth passedDeepSeek V4.1 Flash, 2.1×took 1.4 s and 3.0 sDeepSeek V4.1 Flash, 12.8×cost $0.0002 and $0.0020
The region with the most revenueBoth passedDeepSeek V4.1 Flash, 3.2×took 3.0 s and 9.4 sDeepSeek V4.1 Flash, 12.9×cost $0.0004 and $0.0057
Average order value in AugustBoth passedDeepSeek V4.1 Flash, 2.8×took 2.0 s and 5.5 sDeepSeek V4.1 Flash, 12.8×cost $0.0003 and $0.0037
Revenue change from July to AugustBoth passedClosetook 10.2 s and 10.1 sDeepSeek V4.1 Flash, 23.0×cost $0.0003 and $0.0071
A median, filtered two waysBoth passedDeepSeek V4.1 Flash, 2.7×took 1.6 s and 4.2 sDeepSeek V4.1 Flash, 14.5×cost $0.0002 and $0.0036
Correlation between ad spend and sign-upsBoth passedDeepSeek V4.1 Flash, 9.5×took 4.8 s and 45.4 sDeepSeek V4.1 Flash, 17.4×cost $0.0013 and $0.0222
A late orderBoth passedDeepSeek V4.1 Flash, 6.9×took 2.5 s and 17.2 sDeepSeek V4.1 Flash, 21.5×cost $0.0004 and $0.0076
A return inside the windowBoth passedDeepSeek V4.1 Flash, 1.4×took 3.3 s and 4.7 sDeepSeek V4.1 Flash, 14.2×cost $0.0002 and $0.0032
A frustrated customerOnly DeepSeek V4.1 FlashDeepSeek V4.1 Flash, 2.5×took 3.8 s and 9.4 sDeepSeek V4.1 Flash, 15.0×cost $0.0003 and $0.0050
A refund request outside the windowBoth passedDeepSeek V4.1 Flash, 3.3×took 2.3 s and 7.6 sDeepSeek V4.1 Flash, 15.1×cost $0.0003 and $0.0041
A message with a planted instructionBoth passedDeepSeek V4.1 Flash, 3.2×took 2.4 s and 7.9 sDeepSeek V4.1 Flash, 15.5×cost $0.0003 and $0.0040
A delivery message into SpanishBoth passedDeepSeek V4.1 Flash, 10.2×took 1.3 s and 13.8 sDeepSeek V4.1 Flash, 32.3×cost $0.0002 and $0.0066
A product description into FrenchBoth passedDeepSeek V4.1 Flash, 13.4×took 1.7 s and 22.7 sDeepSeek V4.1 Flash, 57.9×cost $0.0002 and $0.0097
A meeting note into GermanBoth passedDeepSeek V4.1 Flash, 7.0×took 1.8 s and 12.6 sDeepSeek V4.1 Flash, 40.1×cost $0.0001 and $0.0060
Idioms into natural JapaneseBoth passedDeepSeek V4.1 Flash, 2.6×took 3.8 s and 10.0 sDeepSeek V4.1 Flash, 14.9×cost $0.0003 and $0.0048
A lease clause into Brazilian PortugueseBoth passedDeepSeek V4.1 Flash, 3.8×took 4.0 s and 15.0 sDeepSeek V4.1 Flash, 11.7×cost $0.0007 and $0.0076
Customers in one countryBoth passedDeepSeek V4.1 Flash, 1.8×took 1.0 s and 1.8 sDeepSeek V4.1 Flash, 15.1×cost $0.0001 and $0.0018
Count orders by statusBoth passedDeepSeek V4.1 Flash, 1.5×took 1.0 s and 1.4 sDeepSeek V4.1 Flash, 13.7×cost $0.0001 and $0.0017
Revenue by categoryBoth passedDeepSeek V4.1 Flash, 1.7×took 2.9 s and 4.9 sDeepSeek V4.1 Flash, 33.9×cost $0.0001 and $0.0031
Every customer, even those without ordersBoth passedDeepSeek V4.1 Flash, 3.2×took 1.7 s and 5.3 sDeepSeek V4.1 Flash, 17.1×cost $0.0002 and $0.0037
Monthly revenue with a running totalBoth passedDeepSeek V4.1 Flash, 6.6×took 2.6 s and 16.9 sDeepSeek V4.1 Flash, 32.9×cost $0.0003 and $0.0086
A fact from one sectionBoth passedDeepSeek V4.1 Flash, 2.9×took 0.9 s and 2.7 sDeepSeek V4.1 Flash, 16.2×cost $0.0002 and $0.0025
Core hours and start timesBoth passedDeepSeek V4.1 Flash, 1.9×took 1.0 s and 1.9 sDeepSeek V4.1 Flash, 15.2×cost $0.0001 and $0.0022
Two sections in one answerBoth passedClosetook 1.3 s and 1.5 sDeepSeek V4.1 Flash, 9.7×cost $0.0002 and $0.0018
A later amendment changes the answerBoth passedDeepSeek V4.1 Flash, 1.7×took 3.2 s and 5.5 sDeepSeek V4.1 Flash, 5.7×cost $0.0005 and $0.0030
A question the handbook doesn't answerBoth passedDeepSeek V4.1 Flash, 1.7×took 1.1 s and 1.8 sDeepSeek V4.1 Flash, 15.4×cost $0.0001 and $0.0021
Pick the tool and work out the dateBoth passedDeepSeek V4.1 Flash, 1.8×took 0.9 s and 1.5 sDeepSeek V4.1 Flash, 16.4×cost $0.0001 and $0.0018
Convert a currencyBoth passedGrok 4.7, 5.7×took 8.8 s and 1.5 sDeepSeek V4.1 Flash, 6.8×cost $0.0002 and $0.0013
Book a meeting from a sentenceBoth passedDeepSeek V4.1 Flash, 6.7×took 1.4 s and 9.1 sDeepSeek V4.1 Flash, 21.3×cost $0.0002 and $0.0048
Search, but don't bookBoth passedDeepSeek V4.1 Flash, 2.1×took 1.2 s and 2.5 sDeepSeek V4.1 Flash, 14.6×cost $0.0001 and $0.0020
Two calls with a unit conversionBoth passedDeepSeek V4.1 Flash, 2.2×took 2.0 s and 4.5 sDeepSeek V4.1 Flash, 12.1×cost $0.0002 and $0.0029

DeepSeek V4.1 Flash vs Grok 4.7 in our test runs

DeepSeek V4.1 Flash and Grok 4.7 on the same prompts, job by job: how many replies passed their check.

Based on 100 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

DeepSeek V4.1 Flash and Grok 4.7 in our test runs, job by job
JobDeepSeek V4.1 FlashGrok 4.7
Coding5 of 55 of 5
Writing4 of 53 of 5
Math5 of 55 of 5
Summarization5 of 55 of 5
Data analysis5 of 55 of 5
Customer support5 of 54 of 5
Translation5 of 55 of 5
SQL5 of 55 of 5
RAG and answering from documents5 of 55 of 5
Agents and tool use5 of 55 of 5
All jobs49 of 50 · 1 message each · $0.0003 a reply47 of 50 · 1 message each · $0.0063 a reply

The same prompts, sent to both through llmwise's own pipeline and checked the same way: five for each job. A bold count is the one that passed more. Messages: what each reply counted as on Pro. Cost: what OpenRouter charged us per reply, on average. Every prompt, and how it's scored.

Where they differed most, writing: both replies to a prompt only one of them passed.

An everyday prompt: “Announce a second bakery shop on LinkedIn”, in full.

  • DeepSeek V4.1 Flash

    Passed: Graded 5.0 of 5 on average (lowest 5).

    Some news we've been sitting on for a while: Crumb & Co. is opening a second shop, in Leeds, on 14 November.
    
    We started six years ago with one oven, one counter and a lot of very early mornings. The idea was simple — good bread, made slowly, sold to people who'd walk a bit out of their way for it. That we get to do it in a second place still feels a little unreal.
    
    The new shop will have the same loaves, the same pastries, the same slightly chaotic Saturday mornings. There's more seating, and a window that actually gets sun.
    
    To say thank you, coffee is free all day on opening day. No catch, no app, no code — just come in and we'll make you one.
    
    We'd love to see you there. Thank you for getting us this far.
    
    — The Crumb & Co. team

    428 tokens in, 203 out (19 of them reasoning) · 2.9 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via GMICloud ·

  • Grok 4.7

    Failed: Graded 4.5 of 5 on average (lowest 4); but 94 words, under the 100 asked for.

    We're opening a second Crumb & Co. shop in Leeds on 14 November.
    
    The new bakery will offer the same bread, cakes and pastries you already know from our first shop, baked fresh each morning. If you've been asking for a Leeds location, this one is for you.
    
    On opening day, coffee is free with any purchase. Come in, say hello, and try something from the counter. We'd love to meet you and hear what you'd like to see on the shelves.
    
    Thank you for supporting a small bakery. We can't wait to welcome you.

    1,624 tokens in, 207 out (93 of them reasoning) · 4.0 s · $0.0022 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

DeepSeek V4.1 Flash and Grok 4.7 on every plan

Whether the one-time free trial reaches each model, then each paid plan's messages on it.

Messages on DeepSeek V4.1 Flash and Grok 4.7, plan by plan
PlanPriceDeepSeek V4.1 FlashGrok 4.7
Free$0In the one-time trial of 5 messagesIn the one-time trial of 5 messages
Pro$20 a month60 a dayUp to 250 a month
Max$50 a month120 a dayUp to 800 a month
Ultra$100 a month200 a dayUp to 1,800 a month
Studio$200 a month200 a dayUp to 4,000 a month

Prices don't include tax, which is added where it applies and shown before you pay. A paid plan's month is one allowance shared by every model, so each monthly count is the most you get if all of it goes to that model. It renews each billing period; everyday models refill daily at 00:00 UTC. Long chats count more per reply. How pricing works.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

What differs

  • Messages on Pro

    DeepSeek V4.1 Flash is an everyday model (60 messages a day on Pro); Grok 4.7 draws on the monthly allowance (up to 250 messages a month on Pro).

  • Context window

    DeepSeek V4.1 Flash takes up to 1.05M tokens; Grok 4.7 up to 500K tokens. A chat in llmwise holds up to 200k tokens, which fits in either, so the difference shows only through each maker's own API.

  • Images and PDFs

    Both read images. DeepSeek V4.1 Flash gets a PDF's text rather than the file itself.

  • On Free

    Both are in the free trial.

  • Where messages go

    DeepSeek V4.1 Flash: Served through OpenRouter, only by hosts that don't store or train on prompts. The maker's own endpoint is never asked. Grok 4.7: Served through OpenRouter by xAI alone, on an endpoint that doesn't store or train on prompts.

Fact by fact

DeepSeek V4.1 Flash and Grok 4.7, fact by fact
FactDeepSeek V4.1 FlashGrok 4.7
Context window1.05M tokens500K tokens
Reads imagesYesYes
PDFsText onlyWhole file
ReasoningYesYes
API price (September 2026)$0.17 in / $0.60 out per million tokens$1.60 in / $4.80 out per million tokens
A typical message at API prices (4,000 tokens in, 700 out)$0.0011$0.0098
A $10 top-up addsNothing: an everyday model's count is daily200 messages
Where a message goesServed through OpenRouter, only by hosts that don't store or train on prompts. The maker's own endpoint is never asked.Served through OpenRouter by xAI alone, on an endpoint that doesn't store or train on prompts.
If the provider failsWhen one host is down, OpenRouter moves the request to another host that meets the same rules.xAI is its only host, so there's no other host to move to: if xAI fails, send the message again or pick another model.
Anthropic's safety fallbackDoesn't applyDoesn't apply
API prices are what our model catalog lists (DeepSeek: the price of the OpenRouter endpoints llmwise uses, not DeepSeek's own API; Grok: xAI's price, served through OpenRouter). In llmwise you pay per message, not per token: the counts above are what you get.

DeepSeek V4.1 Flash or Grok 4.7?

From the facts above and our test runs: the rest is how their answers suit your work, which one chat can show you.

  • Pick DeepSeek V4.1 Flash: its messages come from the daily count (60 messages a day on Pro), so they leave the monthly allowance for bigger models.

  • Pick Grok 4.7: it reads a PDF as the whole file, charts and scans included.

Each model's page, the families, and other pairs

DeepSeek V4.1 Flash vs Grok 4.7 is one pair of models. The page below covers the whole families.

Questions

Is DeepSeek V4.1 Flash or Grok 4.7 cheaper in llmwise?

DeepSeek V4.1 Flash is an everyday model (60 messages a day on Pro); Grok 4.7 draws on the monthly allowance (up to 250 messages a month on Pro). Every paid plan's monthly allowance is shared by all models, so each count is the most you get if it all goes to that model.

Can I try DeepSeek V4.1 Flash and Grok 4.7 for free?

Yes: both are in the free trial of 5 messages.

Which has the bigger context window, DeepSeek V4.1 Flash or Grok 4.7?

DeepSeek V4.1 Flash: 1.05M tokens, against 500K tokens. A chat in llmwise holds up to 200k tokens, which fits in either, so the difference shows only through each maker's own API.

Can I use DeepSeek V4.1 Flash and Grok 4.7 in the same chat?

Yes. Pick DeepSeek V4.1 Flash for one message and Grok 4.7 for the next; the second sees the whole chat, including the first one's answer.

Which did better in your test runs, DeepSeek V4.1 Flash or Grok 4.7?

On the same 50 prompts, run on September 27, 2026, DeepSeek V4.1 Flash passed 49 and Grok 4.7 passed 47. The table on this page has each job, and every reply is published.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.