Skip to content

Comparison

DeepSeek vs Gemini

In our test runs on October 8, 2026, the same 50 prompts across 10 jobs: DeepSeek's 2 models passed 97 of 100 replies and Gemini's 2 models passed 95 of 100 replies. On llmwise Pro, DeepSeek V4 Pro up to 250 messages a month and Gemini 3.1 Pro (preview) up to 125.

Based on 200 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

Job by job, both families' best models shared the top result on 8 of the 10 jobs; DeepSeek's alone had it on writing; Gemini's alone had it on customer support.

DeepSeek vs Gemini, job by job

On each job, DeepSeek's pick against Gemini's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.

  • Writing: DeepSeek V4 Pro passed 5 of 5 and Gemini 3.1 Pro (preview) 4 of 5. DeepSeek V4 Pro answered 2.3× sooner at the median, 3.9 s against 9.1 s. DeepSeek V4 Pro cost 8.6× less, $0.0054 against $0.0459 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 26% longer, in tokens of reply, thinking not counted.

  • Customer support: DeepSeek V4 Pro passed 4 of 5 and Gemini 3.1 Pro (preview) 5 of 5. DeepSeek V4 Pro answered 2.1× sooner at the median, 4.1 s against 8.7 s. DeepSeek V4 Pro cost 6.9× less, $0.0077 against $0.0529 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 22% longer, in tokens of reply, thinking not counted.

  • Coding: DeepSeek V4.1 Flash passed 5 of 5 and Gemini 3.8 Flash 5 of 5. DeepSeek V4.1 Flash answered 1.7× sooner at the median, 2.5 s against 4.2 s. They cost about the same, $0.0092 against $0.0098 for the 5 replies. Their replies ran to about the same length.

  • Math: DeepSeek V4.1 Flash passed 5 of 5 and Gemini 3.8 Flash 5 of 5. DeepSeek V4.1 Flash answered 5.0× sooner at the median, 0.9 s against 4.7 s. DeepSeek V4.1 Flash cost 5.7× less, $0.0012 against $0.0071 for the 5 replies. Gemini 3.8 Flash's replies ran 70% longer, in tokens of reply, thinking not counted.

  • Summarization: DeepSeek V4.1 Flash passed 5 of 5 and Gemini 3.8 Flash 5 of 5. DeepSeek V4.1 Flash answered 7.2× sooner at the median, 0.5 s against 3.7 s. DeepSeek V4.1 Flash cost 3.7× less, $0.0010 against $0.0038 for the 5 replies. Their replies ran to about the same length.

  • Data analysis: DeepSeek V4.1 Flash passed 5 of 5 and Gemini 3.1 Pro (preview) 5 of 5. DeepSeek V4.1 Flash answered 10.0× sooner at the median, 0.9 s against 8.6 s. DeepSeek V4.1 Flash cost 18.2× less, $0.0042 against $0.0773 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 109% longer, in tokens of reply, thinking not counted.

  • Translation: DeepSeek V4.1 Flash passed 5 of 5 and Gemini 3.8 Flash 5 of 5. DeepSeek V4.1 Flash answered 2.2× sooner at the median, 1.3 s against 3.0 s. DeepSeek V4.1 Flash cost 1.3× less, $0.0029 against $0.0037 for the 5 replies. DeepSeek V4.1 Flash's replies ran 51% longer, in tokens of reply, thinking not counted.

  • SQL: DeepSeek V4.1 Flash passed 5 of 5 and Gemini 3.8 Flash 5 of 5. DeepSeek V4.1 Flash answered 9.0× sooner at the median, 0.4 s against 3.7 s. DeepSeek V4.1 Flash cost 3.7× less, $0.0010 against $0.0037 for the 5 replies. Gemini 3.8 Flash's replies ran 26% longer, in tokens of reply, thinking not counted.

  • RAG and answering from documents: DeepSeek V4.1 Flash passed 5 of 5 and Gemini 3.8 Flash 5 of 5. DeepSeek V4.1 Flash answered 9.2× sooner at the median, 0.3 s against 3.1 s. DeepSeek V4.1 Flash cost 4.0× less, $0.0008 against $0.0033 for the 5 replies. DeepSeek V4.1 Flash's replies ran 57% longer, in tokens of reply, thinking not counted.

  • Agents and tool use: DeepSeek V4.1 Flash passed 5 of 5 and Gemini 3.8 Flash 5 of 5. DeepSeek V4.1 Flash answered 12.4× sooner at the median, 0.4 s against 4.4 s. DeepSeek V4.1 Flash cost 4.3× less, $0.0010 against $0.0043 for the 5 replies. Their replies ran to about the same length.

The 2 prompts only one of DeepSeek and Gemini passed

Where one family's pick passed a prompt and the other's didn't, in each check's own words.

  • Argue both sides of free buses (writing): DeepSeek V4 Pro passed and Gemini 3.1 Pro (preview) didn't. DeepSeek V4 Pro: Graded 4.7 of 5 on average (lowest 4). Gemini 3.1 Pro (preview): Graded 3.7 of 5 on average (lowest 3).

  • A frustrated customer (customer support): Gemini 3.1 Pro (preview) passed and DeepSeek V4 Pro didn't. DeepSeek V4 Pro: Graded 2.3 of 5 on average (lowest 1). Gemini 3.1 Pro (preview): Graded 4.0 of 5 on average (lowest 3).

Each DeepSeek model against each Gemini model

Every DeepSeek model against every Gemini model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.

  • DeepSeek V4 Pro vs Gemini 3.1 Pro: 49 and 49 of 50; 2 prompts split them; DeepSeek V4 Pro's replies cost 2.9× less in all, and DeepSeek V4 Pro answered sooner on 41, Gemini 3.1 Pro on 7. Gemini 3.1 Pro (preview) vs DeepSeek V4 Pro.

  • DeepSeek V4 Pro vs Gemini 3.8 Flash: 49 and 46 of 50; 3 prompts split them; Gemini 3.8 Flash's replies cost 2.8× less in all, and DeepSeek V4 Pro answered sooner on 28, Gemini 3.8 Flash on 17. Gemini 3.8 Flash vs DeepSeek V4 Pro.

  • DeepSeek V4.1 Flash vs Gemini 3.1 Pro: 48 and 49 of 50; 3 prompts split them; DeepSeek V4.1 Flash's replies cost 20.5× less in all, and DeepSeek V4.1 Flash answered sooner on 50, Gemini 3.1 Pro on 0.

  • DeepSeek V4.1 Flash vs Gemini 3.8 Flash: 48 and 46 of 50; 4 prompts split them; DeepSeek V4.1 Flash's replies cost 2.5× less in all, and DeepSeek V4.1 Flash answered sooner on 45, Gemini 3.8 Flash on 1. Gemini 3.8 Flash vs DeepSeek V4.1 Flash.

One writing prompt, both replies

DeepSeek V4 Pro and Gemini 3.1 Pro (preview) on a writing prompt, the job where they differed most: both replies as they came.

A hard prompt: “Argue both sides of free buses”, in full.

  • DeepSeek V4 Pro

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Making buses free turns public transit into a true public good, like sidewalks or libraries. It removes fares as a barrier for low-income riders, boosts ridership, cuts traffic congestion and emissions, and speeds boarding by eliminating payment delays. Cities already subsidize most transit costs; dropping the small fare share is a modest investment with outsized returns in mobility, equity, and cleaner air.
    
    The strongest case against free buses is that fares, however small, are the only direct signal of demand. Removing them invites overcrowding on popular routes while doing nothing to expand service where it is sparse. Scarce transit dollars are better spent on frequency, reliability, and coverage—improvements that actually attract riders—than on subsidizing trips for people who can already afford to pay.

    395 tokens in, 193 out (40 of them reasoning) · 4.1 s · $0.0009 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Wafer ·

  • Gemini 3.1 Pro

    Failed: Graded 3.7 of 5 on average (lowest 3).

    Making city buses free is a powerful catalyst for economic equality and environmental sustainability. By eliminating fares, cities immediately remove a significant financial burden from low-income residents, improving their access to jobs, education, and healthcare. Furthermore, free transit incentivizes people to leave their cars at home, dramatically reducing traffic congestion and cutting harmful greenhouse gas emissions. This shift not only improves urban air quality but also reduces the long-term wear and tear on city infrastructure, ultimately creating a cleaner, fairer, and more accessible environment for all residents.
    
    However, eliminating fares fundamentally undermines the financial stability and quality of the transit system itself. Fares provide crucial revenue required for essential maintenance, route expansion, and driver salaries. Without this income, cities are forced to either slash service frequencies and coverage or drastically raise local taxes to cover the deficit. Consequently, free transit often results in overcrowded, unreliable, and dilapidated buses. Passengers ultimately prioritize speed, safety, and reliability over zero cost; a free bus is useless to a commuter if it never arrives on time.

    409 tokens in, 699 out (492 of them reasoning) · 9.1 s · $0.0092 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

Every model, every job

All 4 DeepSeek and Gemini models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.

Every DeepSeek and Gemini model across our test runs
ModelPassedHard onesCost per replyOn Pro
DeepSeek V4 ProDeepSeek49 of 5020 of 20$0.0037Up to 250 a month
DeepSeek V4.1 FlashDeepSeek48 of 5019 of 20$0.00052Up to 60 a day
Gemini 3.1 Pro (preview)Google49 of 5019 of 20$0.0107Up to 125 a month
Gemini 3.8 FlashGoogle46 of 5018 of 20$0.0013Up to 250 a month

Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Its own subscription

The one company here with its own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.

llmwise Pro, $20 a month, has all 4 of these models in one chat, on one monthly allowance. On it: DeepSeek V4 Pro up to 250 messages a month and Gemini 3.1 Pro (preview) up to 125.

The lineups at a glance

What follows from each model's facts in our catalog.

  • The lineups

    DeepSeek: 2 models, DeepSeek V4 Pro and DeepSeek V4.1 Flash. Gemini: 2 models, Gemini 3.1 Pro (preview) and Gemini 3.8 Flash.

  • Price per message

    The least expensive DeepSeek model is DeepSeek V4.1 Flash (60 messages a day on Pro); the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro).

  • Context window

    Both go up to 1.05M tokens of context (DeepSeek V4 Pro and Gemini 3.1 Pro (preview)).

  • Images and PDFs

    DeepSeek V4 Pro doesn't read images. DeepSeek V4 Pro and DeepSeek V4.1 Flash get a PDF's text rather than the file itself.

  • On the Free plan

    Free's one-time trial of 5 messages covers DeepSeek V4 Pro, DeepSeek V4.1 Flash, Gemini 3.1 Pro (preview), and Gemini 3.8 Flash. Paid plans have every model, with messages every month.

Model by model

Two named models side by side, prompt by prompt, each with its messages on every plan.

More head-to-heads

Each of DeepSeek and Gemini against the other families, every job from the same test runs.

Where your messages go

In llmwise, a message to Gemini goes to its maker, Google, or through OpenRouter when llmwise can't reach the maker directly. DeepSeek models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.

Bar chart: Prompts passed in our test runs, all 10 jobs. DeepSeek V4 Pro: 49 of 50; Gemini 3.1 Pro (preview): 49 of 50; DeepSeek V4.1 Flash: 48 of 50; Gemini 3.8 Flash: 46 of 50.
Our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026: the same prompts for every model, each reply checked the same way.

Questions

Which is better, DeepSeek or Gemini?

In our test runs on October 8, 2026, the same 50 prompts across 10 jobs: DeepSeek's 2 models passed 97 of 100 replies and Gemini's 2 models passed 95 of 100 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; DeepSeek's alone had it on writing; Gemini's alone had it on customer support.

Which is cheaper, DeepSeek or Gemini?

In llmwise, the least expensive DeepSeek model is DeepSeek V4.1 Flash (60 messages a day on Pro), and the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro). At API list prices (October 2026), a typical message of 4,000 tokens in and 700 out costs $0.0020 on DeepSeek V4.1 Flash and $0.0056 on Gemini 3.8 Flash.

Can I use DeepSeek and Gemini in the same chat?

Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.