Skip to content

Comparison

ChatGPT vs Kimi

On llmwise Pro, GPT-6 Sol up to 125 messages a month and Kimi K3 up to 125. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself. We ran the same 50 prompts across 10 jobs on every GPT and Kimi model and published every reply: each family's pick against the other's job by job, the prompts where they split, then their plans and how the lineups differ.

Based on 200 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: GPT's 3 models passed 140 of 150 replies and Kimi's one model passed 46 of 50 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; GPT's alone had it on writing and summarization.

GPT vs Kimi, job by job

On each job, GPT's pick against Kimi's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.

  • Summarization: GPT-6 Astra passed 5 of 5 and Kimi K3 3 of 5. GPT-6 Astra answered 1.2× sooner at the median, 2.9 s against 3.4 s. Kimi K3 cost 3.2× less, $0.0546 against $0.0172 for the 5 replies. Their replies ran to about the same length.

  • Writing: GPT-6 Luna passed 5 of 5 and Kimi K3 4 of 5. GPT-6 Luna answered 1.2× sooner at the median, 2.4 s against 3.0 s. GPT-6 Luna cost 38.5× less, $0.0005 against $0.0192 for the 5 replies. Kimi K3's replies ran 38% longer, in tokens of reply, thinking not counted.

  • Coding: GPT-6 Luna passed 5 of 5 and Kimi K3 5 of 5. Kimi K3 answered 1.1× sooner at the median, 4.6 s against 4.1 s. GPT-6 Luna cost 22.0× less, $0.0014 against $0.0316 for the 5 replies. Kimi K3's replies ran 27% longer, in tokens of reply, thinking not counted.

  • Math: GPT-6 Sol passed 5 of 5 and Kimi K3 5 of 5. GPT-6 Sol answered 2.5× sooner at the median, 2.2 s against 5.4 s. GPT-6 Sol cost 2.1× less, $0.0090 against $0.0187 for the 5 replies. Kimi K3's replies ran 18% longer, in tokens of reply, thinking not counted.

  • Data analysis: GPT-6 Luna passed 5 of 5 and Kimi K3 5 of 5. GPT-6 Luna answered 1.4× sooner at the median, 2.5 s against 3.5 s. GPT-6 Luna cost 38.5× less, $0.0009 against $0.0339 for the 5 replies. Kimi K3's replies ran 129% longer, in tokens of reply, thinking not counted.

  • Customer support: GPT-6 Luna passed 4 of 5 and Kimi K3 4 of 5. GPT-6 Luna answered 3.1× sooner at the median, 1.4 s against 4.3 s. GPT-6 Luna cost 52.7× less, $0.0004 against $0.0228 for the 5 replies. Kimi K3's replies ran 128% longer, in tokens of reply, thinking not counted.

  • Translation: GPT-6 Luna passed 5 of 5 and Kimi K3 5 of 5. GPT-6 Luna answered 3.4× sooner at the median, 1.6 s against 5.2 s. GPT-6 Luna cost 38.6× less, $0.0005 against $0.0179 for the 5 replies. Kimi K3's replies ran 114% longer, in tokens of reply, thinking not counted.

  • SQL: GPT-6 Luna passed 5 of 5 and Kimi K3 5 of 5. GPT-6 Luna answered 1.3× sooner at the median, 1.2 s against 1.6 s. GPT-6 Luna cost 26.6× less, $0.0004 against $0.0118 for the 5 replies. Their replies ran to about the same length.

  • RAG and answering from documents: GPT-6 Luna passed 5 of 5 and Kimi K3 5 of 5. GPT-6 Luna answered 2.1× sooner at the median, 1.3 s against 2.6 s. GPT-6 Luna cost 18.9× less, $0.0004 against $0.0073 for the 5 replies. Kimi K3's replies ran 60% longer, in tokens of reply, thinking not counted.

  • Agents and tool use: GPT-6 Luna passed 5 of 5 and Kimi K3 5 of 5. Their median waits were close, 1.9 s against 2.0 s. GPT-6 Luna cost 33.3× less, $0.0004 against $0.0135 for the 5 replies. Kimi K3's replies ran 34% longer, in tokens of reply, thinking not counted.

The 5 prompts only one of GPT and Kimi passed

Where one family's pick passed a prompt and the other's didn't, in each check's own words.

  • An article in three bullets (summarization): GPT-6 Astra passed and Kimi K3 didn't. GPT-6 Astra: Graded 4.7 of 5 on average (lowest 4). Kimi K3: Graded 4.0 of 5 on average (lowest 3); but 62 words, over the 60 allowed.

  • An email thread in one sentence (summarization): GPT-6 Astra passed and Kimi K3 didn't. GPT-6 Astra: Graded 5.0 of 5 on average (lowest 5). Kimi K3: Graded 5.0 of 5 on average (lowest 5); but 32 words, over the 30 allowed.

  • A product announcement with five rules (writing): GPT-6 Luna passed and Kimi K3 didn't. GPT-6 Luna: Graded 4.0 of 5 on average (lowest 3). Kimi K3: Graded 4.3 of 5 on average (lowest 4); but doesn't end with a question.

  • A frustrated customer (customer support): Kimi K3 passed and GPT-6 Luna didn't. GPT-6 Luna: Graded 3.0 of 5 on average (lowest 2). Kimi K3: Graded 4.0 of 5 on average (lowest 3).

  • A refund request outside the window (customer support): GPT-6 Luna passed and Kimi K3 didn't. GPT-6 Luna: Graded 4.7 of 5 on average (lowest 4). Kimi K3: Graded 4.0 of 5 on average (lowest 2).

Each GPT model against each Kimi model

Every GPT model against every Kimi model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.

  • GPT-6 Astra vs Kimi K3: 48 and 46 of 50; 6 prompts split them; Kimi K3's replies cost 3.0× less in all, and Kimi K3 answered sooner on 23, GPT-6 Astra on 22.

  • GPT-6 Sol vs Kimi K3: 45 and 46 of 50; 7 prompts split them; GPT-6 Sol's replies cost 1.5× less in all, and GPT-6 Sol answered sooner on 32, Kimi K3 on 11. GPT-6 Sol vs Kimi K3.

  • GPT-6 Luna vs Kimi K3: 47 and 46 of 50; 5 prompts split them; GPT-6 Luna's replies cost 32.8× less in all, and GPT-6 Luna answered sooner on 42, Kimi K3 on 5.

One summarization prompt, both replies

GPT-6 Astra and Kimi K3 on a summarization prompt, the job where they differed most: both replies as they came.

An everyday prompt: “An article in three bullets”, in full.

  • GPT-6 Astra

    Passed: Graded 4.7 of 5 on average (lowest 4).

    - Aldmoor’s six-month bike-lane trial ended; November’s council vote decides permanence, costing another €1.1 million after €410,000 installation.
    - Daily bike trips rose from 900 to 2,300; car traffic fell 14%, rush-hour driving times increased just over a minute, and cyclist injuries fell from nine to two.
    - Of 60 businesses, 27 gained customers, 19 lost customers, and 14 saw no change.

    522 tokens in, 182 out (81 of them reasoning) · 5.0 s · $0.0143 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

  • Kimi K3

    Failed: Graded 4.0 of 5 on average (lowest 3); but 62 words, over the 60 allowed.

    - Aldmoor's six-month protected bike lane trial on Market Street ended; the council votes in November on making them permanent.
    - Bike trips rose from 900 to 2,300 daily, cyclist injuries dropped from nine to two, while car traffic fell 14% with only slightly longer rush-hour drive times.
    - Businesses were split on customer impact; installation cost €410,000, with permanent lanes requiring €1.1 million more.

    608 tokens in, 109 out (10 of them reasoning) · 2.8 s · $0.0028 · 1 message on Pro · answered by moonshotai/kimi-k3 via Phala ·

Every model, every job

All 4 GPT and Kimi models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.

Every GPT and Kimi model across our test runs
ModelPassedHard onesCost per replyOn Pro
GPT-6 AstraOpenAI48 of 5020 of 20$0.0117Up to 31 a month
GPT-6 SolOpenAI45 of 5019 of 20$0.0026Up to 125 a month
GPT-6 LunaOpenAI47 of 5020 of 20$0.00012Up to 60 a day
Kimi K3Moonshot46 of 5018 of 20$0.0039Up to 125 a month

Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Its own subscription

The one company here with its own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.

llmwise Pro, $20 a month, has all 4 of these models in one chat, from one monthly allowance: GPT-6 Sol up to 125 messages a month and Kimi K3 up to 125.

The lineups at a glance

What follows from each model's facts in our catalog.

  • The lineups

    GPT: 3 models, GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. Kimi: one model, Kimi K3.

  • Price per message

    The least expensive GPT model is GPT-6 Luna (60 messages a day on Pro); Kimi's one model is Kimi K3 (125 messages a month on Pro).

  • Context window

    GPT goes up to 1.05M tokens (GPT-6 Astra); Kimi up to 1.05M tokens (Kimi K3).

  • Images and PDFs

    Every model here reads images. Kimi K3 gets a PDF's text rather than the file itself.

  • On the Free plan

    Free's one-time trial of 5 messages covers GPT-6 Sol, GPT-6 Luna, and Kimi K3. Paid plans have every model, with messages every month.

Model by model

Two named models side by side, prompt by prompt, each with its messages on every plan.

More head-to-heads

Each of GPT and Kimi against the other families, every job from the same test runs.

Where your messages go

In llmwise, a message to GPT goes to its maker, OpenAI, or through OpenRouter when llmwise can't reach the maker directly. Kimi models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.

Questions

Which is better, ChatGPT or Kimi?

In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: GPT's 3 models passed 140 of 150 replies and Kimi's one model passed 46 of 50 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; GPT's alone had it on writing and summarization.

Which is cheaper, GPT or Kimi?

In llmwise, the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro), and the least expensive Kimi model is Kimi K3 (125 messages a month on Pro). At API list prices (September 2026), a typical message of 4,000 tokens in and 700 out costs $0.0008 on GPT-6 Luna and $0.0225 on Kimi K3.

Can I use GPT and Kimi in the same chat?

Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.