Comparison
Claude vs Kimi
On llmwise Pro, Claude Sonnet 5 up to 125 messages a month and Kimi K3 up to 125. We ran the same 50 prompts across 10 jobs on every Claude and Kimi model and published every reply: each family's pick against the other's job by job, the prompts where they split, then their plans and how the lineups differ.
Based on 250 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: Claude's 4 models passed 183 of 200 replies and Kimi's one model passed 46 of 50 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; Claude's alone had it on summarization and customer support.
Claude vs Kimi, job by job
On each job, Claude's pick against Kimi's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.
Summarization: Claude Haiku 4.5 passed 5 of 5 and Kimi K3 3 of 5. Claude Haiku 4.5 answered 2.0× sooner at the median, 1.7 s against 3.4 s. Claude Haiku 4.5 cost 3.4× less, $0.0051 against $0.0172 for the 5 replies. Kimi K3's replies ran 13% longer, in tokens of reply, thinking not counted.
Customer support: Claude Sonnet 5 passed 5 of 5 and Kimi K3 4 of 5. Claude Sonnet 5 answered 1.2× sooner at the median, 3.5 s against 4.3 s. Claude Sonnet 5 cost 1.3× less, $0.0182 against $0.0228 for the 5 replies. Claude Sonnet 5's replies ran 40% longer, in tokens of reply, thinking not counted.
Coding: Claude Sonnet 5 passed 5 of 5 and Kimi K3 5 of 5. Claude Sonnet 5 answered 1.5× sooner at the median, 2.7 s against 4.1 s. Kimi K3 cost 1.6× less, $0.0508 against $0.0316 for the 5 replies. Claude Sonnet 5's replies ran 50% longer, in tokens of reply, thinking not counted.
Writing: Claude Sonnet 5 passed 4 of 5 and Kimi K3 4 of 5. Kimi K3 answered 1.5× sooner at the median, 4.5 s against 3.0 s. Claude Sonnet 5 cost 1.1× less, $0.0171 against $0.0192 for the 5 replies. Claude Sonnet 5's replies ran 56% longer, in tokens of reply, thinking not counted.
Math: Claude Haiku 4.5 passed 5 of 5 and Kimi K3 5 of 5. Claude Haiku 4.5 answered 2.4× sooner at the median, 2.2 s against 5.4 s. Claude Haiku 4.5 cost 2.4× less, $0.0077 against $0.0187 for the 5 replies. Claude Haiku 4.5's replies ran 185% longer, in tokens of reply, thinking not counted.
Data analysis: Claude Opus 5.5 passed 5 of 5 and Kimi K3 5 of 5. Kimi K3 answered 1.3× sooner at the median, 4.6 s against 3.5 s. Kimi K3 cost 1.9× less, $0.0654 against $0.0339 for the 5 replies. Claude Opus 5.5's replies ran 25% longer, in tokens of reply, thinking not counted.
Translation: Claude Haiku 4.5 passed 5 of 5 and Kimi K3 5 of 5. Claude Haiku 4.5 answered 2.8× sooner at the median, 1.8 s against 5.2 s. Claude Haiku 4.5 cost 2.4× less, $0.0074 against $0.0179 for the 5 replies. Their replies ran to about the same length.
SQL: Claude Haiku 4.5 passed 5 of 5 and Kimi K3 5 of 5. Claude Haiku 4.5 answered 1.3× sooner at the median, 1.2 s against 1.6 s. Claude Haiku 4.5 cost 2.4× less, $0.0049 against $0.0118 for the 5 replies. Their replies ran to about the same length.
RAG and answering from documents: Claude Haiku 4.5 passed 5 of 5 and Kimi K3 5 of 5. Claude Haiku 4.5 answered 2.2× sooner at the median, 1.2 s against 2.6 s. Claude Haiku 4.5 cost 1.4× less, $0.0050 against $0.0073 for the 5 replies. Claude Haiku 4.5's replies ran 46% longer, in tokens of reply, thinking not counted.
Agents and tool use: Claude Haiku 4.5 passed 5 of 5 and Kimi K3 5 of 5. Claude Haiku 4.5 answered 1.8× sooner at the median, 1.1 s against 2.0 s. Claude Haiku 4.5 cost 2.9× less, $0.0047 against $0.0135 for the 5 replies. Claude Haiku 4.5's replies ran 50% longer, in tokens of reply, thinking not counted.
The 5 prompts only one of Claude and Kimi passed
Where one family's pick passed a prompt and the other's didn't, in each check's own words.
An article in three bullets (summarization): Claude Haiku 4.5 passed and Kimi K3 didn't. Claude Haiku 4.5: Graded 4.0 of 5 on average (lowest 3). Kimi K3: Graded 4.0 of 5 on average (lowest 3); but 62 words, over the 60 allowed.
An email thread in one sentence (summarization): Claude Haiku 4.5 passed and Kimi K3 didn't. Claude Haiku 4.5: Graded 4.7 of 5 on average (lowest 4). Kimi K3: Graded 5.0 of 5 on average (lowest 5); but 32 words, over the 30 allowed.
A refund request outside the window (customer support): Claude Sonnet 5 passed and Kimi K3 didn't. Claude Sonnet 5: Graded 4.7 of 5 on average (lowest 4). Kimi K3: Graded 4.0 of 5 on average (lowest 2).
Announce a second bakery shop on LinkedIn (writing): Kimi K3 passed and Claude Sonnet 5 didn't. Claude Sonnet 5: Graded 3.5 of 5 on average (lowest 3). Kimi K3: Graded 4.5 of 5 on average (lowest 4).
A product announcement with five rules (writing): Claude Sonnet 5 passed and Kimi K3 didn't. Claude Sonnet 5: Graded 4.3 of 5 on average (lowest 4). Kimi K3: Graded 4.3 of 5 on average (lowest 4); but doesn't end with a question.
Each Claude model against each Kimi model
Every Claude model against every Kimi model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.
Claude Fable 5.1 vs Kimi K3: 45 and 46 of 50; 3 prompts split them; Kimi K3's replies cost 5.2× less in all, and Kimi K3 answered sooner on 36, Claude Fable 5.1 on 12. Claude Fable 5.1 vs Kimi K3.
Claude Opus 5.5 vs Kimi K3: 49 and 46 of 50; 5 prompts split them; Kimi K3's replies cost 2.8× less in all, and Kimi K3 answered sooner on 33, Claude Opus 5.5 on 11. Claude Opus 5.5 vs Kimi K3.
Claude Sonnet 5 vs Kimi K3: 46 and 46 of 50; 4 prompts split them; their replies cost about the same in all, and Claude Sonnet 5 answered sooner on 23, Kimi K3 on 22. Claude Sonnet 5 vs Kimi K3.
Claude Haiku 4.5 vs Kimi K3: 43 and 46 of 50; 9 prompts split them; Claude Haiku 4.5's replies cost 2.6× less in all, and Claude Haiku 4.5 answered sooner on 40, Kimi K3 on 2. Claude Haiku 4.5 vs Kimi K3.
One summarization prompt, both replies
Claude Haiku 4.5 and Kimi K3 on a summarization prompt, the job where they differed most: both replies as they came.
An everyday prompt: “An article in three bullets”, in full.
Claude Haiku 4.5
Passed: Graded 4.0 of 5 on average (lowest 3).
• **Bike usage tripled** (900 to 2,300 daily trips); cyclist injuries dropped from 9 to 2 during six-month trial. • **Mixed business impact**: 27 of 60 surveyed shops gained customers; 19 lost customers; car travel times increased ~1 minute. • **Council votes November** on permanent lanes costing €1.1 million after €410,000 installation trial.
561 tokens in, 102 out · 1.8 s · $0.0011 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·
Kimi K3
Failed: Graded 4.0 of 5 on average (lowest 3); but 62 words, over the 60 allowed.
- Aldmoor's six-month protected bike lane trial on Market Street ended; the council votes in November on making them permanent. - Bike trips rose from 900 to 2,300 daily, cyclist injuries dropped from nine to two, while car traffic fell 14% with only slightly longer rush-hour drive times. - Businesses were split on customer impact; installation cost €410,000, with permanent lanes requiring €1.1 million more.
608 tokens in, 109 out (10 of them reasoning) · 2.8 s · $0.0028 · 1 message on Pro · answered by moonshotai/kimi-k3 via Phala ·
Every model, every job
All 5 Claude and Kimi models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.
| Model | Passed | Hard ones | Cost per reply | On Pro |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 45 of 50 | 17 of 20 | $0.0200 | Up to 31 a month |
| Claude Opus 5.5Anthropic | 49 of 50 | 19 of 20 | $0.0107 | Up to 62 a month |
| Claude Sonnet 5Anthropic | 46 of 50 | 19 of 20 | $0.0041 | Up to 125 a month |
| Claude Haiku 4.5Anthropic | 43 of 50 | 15 of 20 | $0.0015 | Up to 250 a month |
| Kimi K3Moonshot | 46 of 50 | 18 of 20 | $0.0039 | Up to 125 a month |
Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
Its own subscription
The one company here with its own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.
Claude Pro (Anthropic), $20 a month
Models it names: Claude Opus, Claude Sonnet, and Claude Haiku. On its limits: “The number of messages you can send will vary based on message length, including the length of files you attach, the length of your current conversation, and the model or feature you use. Your session-based usage limit will reset every five hours.” Claude Pro vs llmwise.
Sources: Claude Help Center: What is the Pro plan?, Claude Help Center: Claude Fable models on your plan and Claude pricing, checked .
llmwise Pro, $20 a month, has all 5 of these models in one chat, from one monthly allowance: Claude Sonnet 5 up to 125 messages a month and Kimi K3 up to 125.
The lineups at a glance
What follows from each model's facts in our catalog.
The lineups
Claude: 4 models, Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5, and Claude Haiku 4.5. Kimi: one model, Kimi K3.
Price per message
The least expensive Claude model is Claude Haiku 4.5 (250 messages a month on Pro); Kimi's one model is Kimi K3 (125 messages a month on Pro).
Context window
Claude goes up to 1M tokens (Claude Fable 5.1); Kimi up to 1.05M tokens (Kimi K3).
Images and PDFs
Every model here reads images. Kimi K3 gets a PDF's text rather than the file itself.
On the Free plan
Free's one-time trial of 5 messages covers Claude Sonnet 5, Claude Haiku 4.5, and Kimi K3. Paid plans have every model, with messages every month.
Model by model
Two named models side by side, prompt by prompt, each with its messages on every plan.
More head-to-heads
Each of Claude and Kimi against the other families, every job from the same test runs.
Where your messages go
In llmwise, a message to Claude goes to its maker, Anthropic, or through OpenRouter when llmwise can't reach the maker directly. Kimi models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.
Questions
Which is better, Claude or Kimi?
In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: Claude's 4 models passed 183 of 200 replies and Kimi's one model passed 46 of 50 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; Claude's alone had it on summarization and customer support.
Which is cheaper, Claude or Kimi?
In llmwise, the least expensive Claude model is Claude Haiku 4.5 (250 messages a month on Pro), and the least expensive Kimi model is Kimi K3 (125 messages a month on Pro). At API list prices (September 2026), a typical message of 4,000 tokens in and 700 out costs $0.0075 on Claude Haiku 4.5 and $0.0225 on Kimi K3.
Can I use Claude and Kimi in the same chat?
Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.