Comparison
DeepSeek vs Kimi
On llmwise Pro, DeepSeek V4 Pro up to 250 messages a month and Kimi K3 up to 125. We ran the same 50 prompts across 10 jobs on every DeepSeek and Kimi model and published every reply: each family's pick against the other's job by job, the prompts where they split, then their plans and how the lineups differ.
Based on 150 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: DeepSeek's 2 models passed 94 of 100 replies and Kimi's one model passed 46 of 50 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; DeepSeek's alone had it on summarization and customer support.
DeepSeek vs Kimi, job by job
On each job, DeepSeek's pick against Kimi's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.
Summarization: DeepSeek V4.1 Flash passed 5 of 5 and Kimi K3 3 of 5. DeepSeek V4.1 Flash answered 2.2× sooner at the median, 1.5 s against 3.4 s. DeepSeek V4.1 Flash cost 20.7× less, $0.0008 against $0.0172 for the 5 replies. Kimi K3's replies ran 17% longer, in tokens of reply, thinking not counted.
Customer support: DeepSeek V4.1 Flash passed 5 of 5 and Kimi K3 4 of 5. DeepSeek V4.1 Flash answered 1.7× sooner at the median, 2.5 s against 4.3 s. DeepSeek V4.1 Flash cost 15.8× less, $0.0014 against $0.0228 for the 5 replies. Their replies ran to about the same length.
Coding: DeepSeek V4.1 Flash passed 5 of 5 and Kimi K3 5 of 5. Kimi K3 answered 1.9× sooner at the median, 7.8 s against 4.1 s. DeepSeek V4.1 Flash cost 5.8× less, $0.0054 against $0.0316 for the 5 replies. DeepSeek V4.1 Flash's replies ran 17% longer, in tokens of reply, thinking not counted.
Writing: DeepSeek V4.1 Flash passed 4 of 5 and Kimi K3 4 of 5. Their median waits were close, 3.1 s against 3.0 s. DeepSeek V4.1 Flash cost 13.4× less, $0.0014 against $0.0192 for the 5 replies. Kimi K3's replies ran 18% longer, in tokens of reply, thinking not counted.
Math: DeepSeek V4.1 Flash passed 5 of 5 and Kimi K3 5 of 5. DeepSeek V4.1 Flash answered 3.2× sooner at the median, 1.7 s against 5.4 s. DeepSeek V4.1 Flash cost 16.1× less, $0.0012 against $0.0187 for the 5 replies. DeepSeek V4.1 Flash's replies ran 27% longer, in tokens of reply, thinking not counted.
Data analysis: DeepSeek V4.1 Flash passed 5 of 5 and Kimi K3 5 of 5. DeepSeek V4.1 Flash answered 1.2× sooner at the median, 3.0 s against 3.5 s. DeepSeek V4.1 Flash cost 13.2× less, $0.0026 against $0.0339 for the 5 replies. Their replies ran to about the same length.
Translation: DeepSeek V4.1 Flash passed 5 of 5 and Kimi K3 5 of 5. DeepSeek V4.1 Flash answered 2.9× sooner at the median, 1.8 s against 5.2 s. DeepSeek V4.1 Flash cost 11.9× less, $0.0015 against $0.0179 for the 5 replies. Kimi K3's replies ran 30% longer, in tokens of reply, thinking not counted.
SQL: DeepSeek V4.1 Flash passed 5 of 5 and Kimi K3 5 of 5. Their median waits were close, 1.7 s against 1.6 s. DeepSeek V4.1 Flash cost 14.6× less, $0.0008 against $0.0118 for the 5 replies. Their replies ran to about the same length.
RAG and answering from documents: DeepSeek V4.1 Flash passed 5 of 5 and Kimi K3 5 of 5. DeepSeek V4.1 Flash answered 2.4× sooner at the median, 1.1 s against 2.6 s. DeepSeek V4.1 Flash cost 6.4× less, $0.0011 against $0.0073 for the 5 replies. DeepSeek V4.1 Flash's replies ran 32% longer, in tokens of reply, thinking not counted.
Agents and tool use: DeepSeek V4.1 Flash passed 5 of 5 and Kimi K3 5 of 5. DeepSeek V4.1 Flash answered 1.5× sooner at the median, 1.4 s against 2.0 s. DeepSeek V4.1 Flash cost 14.8× less, $0.0009 against $0.0135 for the 5 replies. Kimi K3's replies ran 13% longer, in tokens of reply, thinking not counted.
The 5 prompts only one of DeepSeek and Kimi passed
Where one family's pick passed a prompt and the other's didn't, in each check's own words.
An article in three bullets (summarization): DeepSeek V4.1 Flash passed and Kimi K3 didn't. DeepSeek V4.1 Flash: Graded 4.7 of 5 on average (lowest 4). Kimi K3: Graded 4.0 of 5 on average (lowest 3); but 62 words, over the 60 allowed.
An email thread in one sentence (summarization): DeepSeek V4.1 Flash passed and Kimi K3 didn't. DeepSeek V4.1 Flash: Graded 4.7 of 5 on average (lowest 4). Kimi K3: Graded 5.0 of 5 on average (lowest 5); but 32 words, over the 30 allowed.
A refund request outside the window (customer support): DeepSeek V4.1 Flash passed and Kimi K3 didn't. DeepSeek V4.1 Flash: Graded 4.7 of 5 on average (lowest 4). Kimi K3: Graded 4.0 of 5 on average (lowest 2).
Rewrite corporate jargon in plain words (writing): Kimi K3 passed and DeepSeek V4.1 Flash didn't. DeepSeek V4.1 Flash: Graded 3.7 of 5 on average (lowest 3). Kimi K3: Graded 4.3 of 5 on average (lowest 4).
A product announcement with five rules (writing): DeepSeek V4.1 Flash passed and Kimi K3 didn't. DeepSeek V4.1 Flash: Graded 4.3 of 5 on average (lowest 4). Kimi K3: Graded 4.3 of 5 on average (lowest 4); but doesn't end with a question.
Each DeepSeek model against each Kimi model
Every DeepSeek model against every Kimi model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.
DeepSeek V4 Pro vs Kimi K3: 45 and 46 of 50; 5 prompts split them; DeepSeek V4 Pro's replies cost 1.9× less in all, and Kimi K3 answered sooner on 25, DeepSeek V4 Pro on 24. DeepSeek V4 Pro vs Kimi K3.
DeepSeek V4.1 Flash vs Kimi K3: 49 and 46 of 50; 5 prompts split them; DeepSeek V4.1 Flash's replies cost 11.3× less in all, and DeepSeek V4.1 Flash answered sooner on 33, Kimi K3 on 10.
One summarization prompt, both replies
DeepSeek V4.1 Flash and Kimi K3 on a summarization prompt, the job where they differed most: both replies as they came.
An everyday prompt: “An article in three bullets”, in full.
DeepSeek V4.1 Flash
Passed: Graded 4.7 of 5 on average (lowest 4).
- Aldmoor's six-month Market Street bike lane trial ended; council votes in November. - Cyclist trips rose 900→2,300 daily, injuries fell 9→2; car traffic fell 14%, rush-hour delays +1 minute. - Businesses split; €410k installed, €1.1m more to make permanent.
562 tokens in, 310 out (238 of them reasoning) · 2.5 s · $0.0003 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via StreamLake ·
Kimi K3
Failed: Graded 4.0 of 5 on average (lowest 3); but 62 words, over the 60 allowed.
- Aldmoor's six-month protected bike lane trial on Market Street ended; the council votes in November on making them permanent. - Bike trips rose from 900 to 2,300 daily, cyclist injuries dropped from nine to two, while car traffic fell 14% with only slightly longer rush-hour drive times. - Businesses were split on customer impact; installation cost €410,000, with permanent lanes requiring €1.1 million more.
608 tokens in, 109 out (10 of them reasoning) · 2.8 s · $0.0028 · 1 message on Pro · answered by moonshotai/kimi-k3 via Phala ·
Every model, every job
All 3 DeepSeek and Kimi models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.
| Model | Passed | Hard ones | Cost per reply | On Pro |
|---|---|---|---|---|
| DeepSeek V4 ProDeepSeek | 45 of 50 | 17 of 20 | $0.0021 | Up to 250 a month |
| DeepSeek V4.1 FlashDeepSeek | 49 of 50 | 20 of 20 | $0.00034 | Up to 60 a day |
| Kimi K3Moonshot | 46 of 50 | 18 of 20 | $0.0039 | Up to 125 a month |
Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
The lineups at a glance
What follows from each model's facts in our catalog.
The lineups
DeepSeek: 2 models, DeepSeek V4 Pro and DeepSeek V4.1 Flash. Kimi: one model, Kimi K3.
Price per message
The least expensive DeepSeek model is DeepSeek V4.1 Flash (60 messages a day on Pro); Kimi's one model is Kimi K3 (125 messages a month on Pro).
Context window
Both go up to 1.05M tokens of context (DeepSeek V4 Pro and Kimi K3).
Images and PDFs
DeepSeek V4 Pro doesn't read images. DeepSeek V4 Pro, DeepSeek V4.1 Flash, and Kimi K3 get a PDF's text rather than the file itself.
On the Free plan
Free's one-time trial of 5 messages covers DeepSeek V4 Pro, DeepSeek V4.1 Flash, and Kimi K3. Paid plans have every model, with messages every month.
Model by model
Two named models side by side, prompt by prompt, each with its messages on every plan.
More head-to-heads
Each of DeepSeek and Kimi against the other families, every job from the same test runs.
Where your messages go
DeepSeek and Kimi models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.
Questions
Which is better, DeepSeek or Kimi?
In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: DeepSeek's 2 models passed 94 of 100 replies and Kimi's one model passed 46 of 50 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; DeepSeek's alone had it on summarization and customer support.
Which is cheaper, DeepSeek or Kimi?
In llmwise, the least expensive DeepSeek model is DeepSeek V4.1 Flash (60 messages a day on Pro), and the least expensive Kimi model is Kimi K3 (125 messages a month on Pro). At API list prices (September 2026), a typical message of 4,000 tokens in and 700 out costs $0.0011 on DeepSeek V4.1 Flash and $0.0225 on Kimi K3.
Can I use DeepSeek and Kimi in the same chat?
Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.