Skip to content

Comparison

Grok vs DeepSeek

On llmwise Pro, Grok 4.7 up to 250 messages a month and DeepSeek V4 Pro up to 250. We ran the same 50 prompts across 10 jobs on every Grok and DeepSeek model and published every reply: each family's pick against the other's job by job, the prompts where they split, then their plans and how the lineups differ.

Based on 150 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: Grok's one model passed 47 of 50 replies and DeepSeek's 2 models passed 94 of 100 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; DeepSeek's alone had it on writing and customer support.

Grok vs DeepSeek, job by job

On each job, Grok's pick against DeepSeek's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.

  • Writing: Grok 4.7 passed 3 of 5 and DeepSeek V4.1 Flash 4 of 5. DeepSeek V4.1 Flash answered 1.3× sooner at the median, 4.0 s against 3.1 s. DeepSeek V4.1 Flash cost 14.7× less, $0.0210 against $0.0014 for the 5 replies. DeepSeek V4.1 Flash's replies ran 31% longer, in tokens of reply, thinking not counted.

  • Customer support: Grok 4.7 passed 4 of 5 and DeepSeek V4.1 Flash 5 of 5. DeepSeek V4.1 Flash answered 3.1× sooner at the median, 7.9 s against 2.5 s. DeepSeek V4.1 Flash cost 16.6× less, $0.0239 against $0.0014 for the 5 replies. DeepSeek V4.1 Flash's replies ran 34% longer, in tokens of reply, thinking not counted.

  • Coding: Grok 4.7 passed 5 of 5 and DeepSeek V4.1 Flash 5 of 5. DeepSeek V4.1 Flash answered 3.3× sooner at the median, 25.3 s against 7.8 s. DeepSeek V4.1 Flash cost 20.3× less, $0.1099 against $0.0054 for the 5 replies. DeepSeek V4.1 Flash's replies ran 24% longer, in tokens of reply, thinking not counted.

  • Math: Grok 4.7 passed 5 of 5 and DeepSeek V4.1 Flash 5 of 5. DeepSeek V4.1 Flash answered 5.4× sooner at the median, 9.1 s against 1.7 s. DeepSeek V4.1 Flash cost 21.7× less, $0.0251 against $0.0012 for the 5 replies. Their replies ran to about the same length.

  • Summarization: Grok 4.7 passed 5 of 5 and DeepSeek V4.1 Flash 5 of 5. DeepSeek V4.1 Flash answered 3.1× sooner at the median, 4.7 s against 1.5 s. DeepSeek V4.1 Flash cost 19.3× less, $0.0161 against $0.0008 for the 5 replies. Their replies ran to about the same length.

  • Data analysis: Grok 4.7 passed 5 of 5 and DeepSeek V4.1 Flash 5 of 5. DeepSeek V4.1 Flash answered 3.2× sooner at the median, 9.4 s against 3.0 s. DeepSeek V4.1 Flash cost 16.5× less, $0.0423 against $0.0026 for the 5 replies. DeepSeek V4.1 Flash's replies ran 55% longer, in tokens of reply, thinking not counted.

  • Translation: Grok 4.7 passed 5 of 5 and DeepSeek V4.1 Flash 5 of 5. DeepSeek V4.1 Flash answered 7.6× sooner at the median, 13.8 s against 1.8 s. DeepSeek V4.1 Flash cost 23.2× less, $0.0347 against $0.0015 for the 5 replies. DeepSeek V4.1 Flash's replies ran 81% longer, in tokens of reply, thinking not counted.

  • SQL: Grok 4.7 passed 5 of 5 and DeepSeek V4.1 Flash 5 of 5. DeepSeek V4.1 Flash answered 3.0× sooner at the median, 4.9 s against 1.7 s. DeepSeek V4.1 Flash cost 23.3× less, $0.0189 against $0.0008 for the 5 replies. DeepSeek V4.1 Flash's replies ran 11% longer, in tokens of reply, thinking not counted.

  • RAG and answering from documents: Grok 4.7 passed 5 of 5 and DeepSeek V4.1 Flash 5 of 5. DeepSeek V4.1 Flash answered 1.7× sooner at the median, 1.9 s against 1.1 s. DeepSeek V4.1 Flash cost 10.1× less, $0.0115 against $0.0011 for the 5 replies. DeepSeek V4.1 Flash's replies ran 160% longer, in tokens of reply, thinking not counted.

  • Agents and tool use: Grok 4.7 passed 5 of 5 and DeepSeek V4.1 Flash 5 of 5. DeepSeek V4.1 Flash answered 1.8× sooner at the median, 2.5 s against 1.4 s. DeepSeek V4.1 Flash cost 14.1× less, $0.0129 against $0.0009 for the 5 replies. DeepSeek V4.1 Flash's replies ran 10% longer, in tokens of reply, thinking not counted.

The 2 prompts only one of Grok and DeepSeek passed

Where one family's pick passed a prompt and the other's didn't, in each check's own words.

  • Announce a second bakery shop on LinkedIn (writing): DeepSeek V4.1 Flash passed and Grok 4.7 didn't. Grok 4.7: Graded 4.5 of 5 on average (lowest 4); but 94 words, under the 100 asked for. DeepSeek V4.1 Flash: Graded 5.0 of 5 on average (lowest 5).

  • A frustrated customer (customer support): DeepSeek V4.1 Flash passed and Grok 4.7 didn't. Grok 4.7: Graded 3.3 of 5 on average (lowest 3). DeepSeek V4.1 Flash: Graded 4.7 of 5 on average (lowest 4).

Each Grok model against each DeepSeek model

Every Grok model against every DeepSeek model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.

  • Grok 4.7 vs DeepSeek V4 Pro: 47 and 45 of 50; 6 prompts split them; DeepSeek V4 Pro's replies cost 3.1× less in all, and DeepSeek V4 Pro answered sooner on 36, Grok 4.7 on 11. DeepSeek V4 Pro vs Grok 4.7.

  • Grok 4.7 vs DeepSeek V4.1 Flash: 47 and 49 of 50; 2 prompts split them; DeepSeek V4.1 Flash's replies cost 18.4× less in all, and DeepSeek V4.1 Flash answered sooner on 45, Grok 4.7 on 3. DeepSeek V4.1 Flash vs Grok 4.7.

One writing prompt, both replies

Grok 4.7 and DeepSeek V4.1 Flash on a writing prompt, the job where they differed most: both replies as they came.

An everyday prompt: “Announce a second bakery shop on LinkedIn”, in full.

  • Grok 4.7

    Failed: Graded 4.5 of 5 on average (lowest 4); but 94 words, under the 100 asked for.

    We're opening a second Crumb & Co. shop in Leeds on 14 November.
    
    The new bakery will offer the same bread, cakes and pastries you already know from our first shop, baked fresh each morning. If you've been asking for a Leeds location, this one is for you.
    
    On opening day, coffee is free with any purchase. Come in, say hello, and try something from the counter. We'd love to meet you and hear what you'd like to see on the shelves.
    
    Thank you for supporting a small bakery. We can't wait to welcome you.

    1,624 tokens in, 207 out (93 of them reasoning) · 4.0 s · $0.0022 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  • DeepSeek V4.1 Flash

    Passed: Graded 5.0 of 5 on average (lowest 5).

    Some news we've been sitting on for a while: Crumb & Co. is opening a second shop, in Leeds, on 14 November.
    
    We started six years ago with one oven, one counter and a lot of very early mornings. The idea was simple — good bread, made slowly, sold to people who'd walk a bit out of their way for it. That we get to do it in a second place still feels a little unreal.
    
    The new shop will have the same loaves, the same pastries, the same slightly chaotic Saturday mornings. There's more seating, and a window that actually gets sun.
    
    To say thank you, coffee is free all day on opening day. No catch, no app, no code — just come in and we'll make you one.
    
    We'd love to see you there. Thank you for getting us this far.
    
    — The Crumb & Co. team

    428 tokens in, 203 out (19 of them reasoning) · 2.9 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via GMICloud ·

Every model, every job

All 3 Grok and DeepSeek models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.

Every Grok and DeepSeek model across our test runs
ModelPassedHard onesCost per replyOn Pro
Grok 4.7xAI47 of 5020 of 20$0.0063Up to 250 a month
DeepSeek V4 ProDeepSeek45 of 5017 of 20$0.0021Up to 250 a month
DeepSeek V4.1 FlashDeepSeek49 of 5020 of 20$0.00034Up to 60 a day

Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Its own subscription

The one company here with its own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.

  • SuperGrok (xAI), $30 a month

    Models it names: Grok 4.6. On its limits: “Smarter answers in Expert mode: Access our best model with higher limits” SuperGrok vs llmwise.

    Source: Grok plans, checked .

llmwise Pro, $20 a month, has all 3 of these models in one chat, from one monthly allowance: Grok 4.7 up to 250 messages a month and DeepSeek V4 Pro up to 250.

The lineups at a glance

What follows from each model's facts in our catalog.

  • The lineups

    Grok: one model, Grok 4.7. DeepSeek: 2 models, DeepSeek V4 Pro and DeepSeek V4.1 Flash.

  • Price per message

    Grok's one model is Grok 4.7 (250 messages a month on Pro); the least expensive DeepSeek model is DeepSeek V4.1 Flash (60 messages a day on Pro).

  • Context window

    Grok goes up to 500K tokens (Grok 4.7); DeepSeek up to 1.05M tokens (DeepSeek V4 Pro).

  • Images and PDFs

    DeepSeek V4 Pro doesn't read images. DeepSeek V4 Pro and DeepSeek V4.1 Flash get a PDF's text rather than the file itself.

  • On the Free plan

    Free's one-time trial of 5 messages covers Grok 4.7, DeepSeek V4 Pro, and DeepSeek V4.1 Flash. Paid plans have every model, with messages every month.

Model by model

Two named models side by side, prompt by prompt, each with its messages on every plan.

More head-to-heads

Each of Grok and DeepSeek against the other families, every job from the same test runs.

Where your messages go

Grok and DeepSeek models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.

Questions

Which is better, Grok or DeepSeek?

In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: Grok's one model passed 47 of 50 replies and DeepSeek's 2 models passed 94 of 100 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; DeepSeek's alone had it on writing and customer support.

Which is cheaper, Grok or DeepSeek?

In llmwise, the least expensive Grok model is Grok 4.7 (250 messages a month on Pro), and the least expensive DeepSeek model is DeepSeek V4.1 Flash (60 messages a day on Pro). At API list prices (September 2026), a typical message of 4,000 tokens in and 700 out costs $0.0098 on Grok 4.7 and $0.0011 on DeepSeek V4.1 Flash.

Can I use Grok and DeepSeek in the same chat?

Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.