Skip to content

Comparison · Math

ChatGPT vs DeepSeek for math

On llmwise Pro, GPT-6 Sol gets up to 125 messages a month and DeepSeek V4 Pro up to 250 messages a month. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself. We ran the same 5 math prompts on all 3 GPT models and all 2 DeepSeek models and published every reply: the results, then GPT-6 Sol against DeepSeek V4.1 Flash prompt by prompt, then how GPT and DeepSeek compare on price per message, context and files.

Based on 25 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

In our math test runs on September 27, 2026, GPT's 3 models passed 14 of 15; GPT-6 Astra and GPT-6 Sol each passed 5 of 5. DeepSeek's 2 models passed 10 of 10; DeepSeek V4 Pro and DeepSeek V4.1 Flash each passed 5 of 5. GPT-6 Sol and DeepSeek V4.1 Flash each passed 5 of the 5 prompts, so these math prompts don't split GPT and DeepSeek; the replies on this page show how they differ.

GPT and DeepSeek on our math test runs

Every GPT and DeepSeek model in llmwise on our 5 math prompts: how many replies passed, what each counted as on Pro, and what it cost to run.

GPT and DeepSeek on our math test runs
ModelPassedHard onesMessages used on ProCost per replyTime per reply
GPT-6 AstraOpenAI5 of 52 of 21 each, of 31 a month on Pro$0.00852.7 s
GPT-6 SolOpenAI5 of 52 of 21 each, of 125 a month on Pro$0.00182.2 s
GPT-6 LunaOpenAI4 of 52 of 21 each, of 60 a day on Pro$0.00012.4 s
DeepSeek V4 ProDeepSeek5 of 52 of 21 each, of 250 a month on Pro$0.00072.7 s
DeepSeek V4.1 FlashDeepSeek5 of 52 of 21 each, of 60 a day on Pro$0.00026.4 s
GPT: 14 of 15 replies passed; DeepSeek: 10 of 10 replies passed. Messages used is what each reply counted as on Pro; cost is what OpenRouter charged us per reply.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

GPT vs DeepSeek, prompt by prompt

Each math prompt with GPT-6 Sol and DeepSeek V4.1 Flash side by side, as they replied: for each side, the model that passed the most math prompts (then the most hard ones, then the cheaper).

Math: GPT-6 Sol passed 5 of 5 and DeepSeek V4.1 Flash 5 of 5. DeepSeek V4.1 Flash answered 1.3× sooner at the median, 2.2 s against 1.7 s. DeepSeek V4.1 Flash cost 7.8× less, $0.0090 against $0.0012 for the 5 replies. DeepSeek V4.1 Flash's replies ran 49% longer, in tokens of reply, thinking not counted.

  1. 1. A discount, then sales tax

    Everyday. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · DeepSeek: 2 of 2 models passed, DeepSeek V4 Pro the cheapest at $0.0002. The prompt in full.

    GPT-6 Sol and DeepSeek V4.1 Flash both passed this one. DeepSeek V4.1 Flash answered sooner (0.6 s to 1.7 s), and DeepSeek V4.1 Flash's reply cost less ($0.0002 to $0.0013); GPT-6 Sol wrote 45 tokens of reply, DeepSeek V4.1 Flash 41.

    • GPT-6 Sol

      Passed: Final answer 66: right.

      25% off $80 gives $80 × 0.75 = $60.  
      Adding 10% sales tax gives $60 × 1.10 = $66.
      
      Final answer: $66

      400 tokens in, 45 out · 1.7 s · $0.0013 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

    • DeepSeek V4.1 Flash

      Passed: Final answer 66: right.

      $80 × 0.75 = $60 (after 25% discount)
      $60 × 1.10 = $66 (with 10% tax)
      
      Final answer: $66

      440 tokens in, 64 out (23 of them reasoning) · 0.6 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Parasail ·

  2. 2. Pens at 3 for $4

    Everyday. GPT: 2 of 3 models passed, GPT-6 Sol the cheapest at $0.0024 · DeepSeek: 2 of 2 models passed, DeepSeek V4.1 Flash the cheapest at $0.0003. The prompt in full.

    GPT-6 Sol and DeepSeek V4.1 Flash both passed this one. GPT-6 Sol answered sooner (3.3 s to 25.6 s), and DeepSeek V4.1 Flash's reply cost less ($0.0003 to $0.0024); GPT-6 Sol wrote 56 tokens of reply, DeepSeek V4.1 Flash 138.

    • GPT-6 Sol

      Passed: Final answer 13.50: right.

      Three groups of 3 pens cost \(3 \times \$4 = \$12\). One single pen costs \$1.50, bringing the total to \$13.50. Using fewer groups would cost more.
      
      Final answer: $13.50

      402 tokens in, 164 out (108 of them reasoning) · 3.3 s · $0.0024 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

    • DeepSeek V4.1 Flash

      Passed: Final answer Buy three 3-packs (9 pens) plus one single pen for 13.50.: right.

      **Working:**
      - 3-pack price per pen: $4 ÷ 3 ≈ $1.33 (best value)
      - Buy as many 3-packs as fit: 3 packs = 9 pens = $12, plus 1 single at $1.50 = **$13.50**
      - Check other combinations: 2 packs + 4 singles = $8 + $6 = $14; 1 pack + 7 singles = $14.50; 10 singles = $15
      - The $13.50 option is cheapest.
      
      Final answer: Buy three 3-packs (9 pens) plus one single pen for $13.50.

      442 tokens in, 854 out (716 of them reasoning) · 25.6 s · $0.0003 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Morph ·

  3. 3. Compound interest over three years

    Everyday. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · DeepSeek: 2 of 2 models passed, DeepSeek V4.1 Flash the cheapest at $0.0002. The prompt in full.

    GPT-6 Sol and DeepSeek V4.1 Flash both passed this one. They took about as long (GPT-6 Sol 1.4 s, DeepSeek V4.1 Flash 1.5 s), and DeepSeek V4.1 Flash's reply cost less ($0.0002 to $0.0014); GPT-6 Sol wrote 62 tokens of reply, DeepSeek V4.1 Flash 104.

    • GPT-6 Sol

      Passed: Final answer 315.25: right.

      After 3 years, the balance is \(2000(1.05)^3=\$2,315.25\). Subtracting the original deposit gives \(2315.25-2000=\$315.25\) in interest.
      
      Final answer: $315.25

      402 tokens in, 62 out · 1.4 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

    • DeepSeek V4.1 Flash

      Passed: Final answer 315.25: right.

      **Working:**
      
      - Year 1: $2,000 × 1.05 = $2,100.00
      - Year 2: $2,100.00 × 1.05 = $2,205.00
      - Year 3: $2,205.00 × 1.05 = $2,315.25
      
      Interest earned = $2,315.25 − $2,000.00 = **$315.25**
      
      Final answer: $315.25

      440 tokens in, 186 out (82 of them reasoning) · 1.5 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via CoreWeave ·

  4. 4. Four-digit numbers whose digits sum to 9

    Hard. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · DeepSeek: 2 of 2 models passed, DeepSeek V4.1 Flash the cheapest at $0.0002. The prompt in full.

    GPT-6 Sol and DeepSeek V4.1 Flash both passed this one. DeepSeek V4.1 Flash answered sooner (1.7 s to 2.5 s), and DeepSeek V4.1 Flash's reply cost less ($0.0002 to $0.0019); GPT-6 Sol wrote 78 tokens of reply, DeepSeek V4.1 Flash 112.

    • GPT-6 Sol

      Passed: Final answer 165: right.

      Subtract 1 from the thousands digit. The four digits then become nonnegative integers that sum to 8. By stars and bars, the number of solutions is \(\binom{8+4-1}{4-1}=\binom{11}{3}=165\). Every solution gives a valid four-digit integer.
      
      Final answer: 165

      378 tokens in, 117 out (39 of them reasoning) · 2.5 s · $0.0019 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

    • DeepSeek V4.1 Flash

      Passed: Final answer 165: right.

      Let the digits be $a,b,c,d$ with $a \ge 1$ (four-digit) and $b,c,d \ge 0$, and $a+b+c+d=9$.
      
      Set $a' = a-1 \ge 0$. Then $a'+b+c+d = 8$.
      
      Number of nonnegative solutions to a sum of 8 across 4 variables:
      
      $$\binom{8+4-1}{4-1} = \binom{11}{3} = 165$$
      
      Final answer: 165

      417 tokens in, 216 out (104 of them reasoning) · 1.7 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Novita ·

  5. 5. The highest of three dice is a 5

    Hard. GPT: 3 of 3 models passed, GPT-6 Luna the cheapest at $0.0001 · DeepSeek: 2 of 2 models passed, DeepSeek V4.1 Flash the cheapest at $0.0002. The prompt in full.

    GPT-6 Sol and DeepSeek V4.1 Flash both passed this one. GPT-6 Sol answered sooner (2.2 s to 2.5 s), and DeepSeek V4.1 Flash's reply cost less ($0.0002 to $0.0020); GPT-6 Sol wrote 91 tokens of reply, DeepSeek V4.1 Flash 100.

    • GPT-6 Sol

      Passed: Final answer 61/216: right.

      For the highest number to be exactly 5, all three dice must show 5 or less, but not all can show 4 or less.
      
      \[
      P(\text{highest is }5)=\left(\frac56\right)^3-\left(\frac46\right)^3
      =\frac{125-64}{216}=\frac{61}{216}.
      \]
      
      Final answer: 61/216

      393 tokens in, 122 out (31 of them reasoning) · 2.2 s · $0.0020 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

    • DeepSeek V4.1 Flash

      Passed: Final answer 61/216: right.

      Total outcomes: 6³ = 216.
      
      Max exactly 5 means all dice show ≤ 5, but not all show ≤ 4:
      
      - All ≤ 5: 5³ = 125
      - All ≤ 4: 4³ = 64
      
      So favorable outcomes: 125 − 64 = 61.
      
      Probability = 61/216 (already in lowest terms, since 61 is prime and doesn't divide 216).
      
      Final answer: 61/216

      432 tokens in, 155 out (55 of them reasoning) · 2.5 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via CoreWeave ·

Each GPT model against each DeepSeek model

Every GPT model against every DeepSeek model on the same 5 math prompts: passes, the median wait and what the replies cost. A lead under 10% counts as close.

  • GPT-6 Astra vs DeepSeek V4 Pro: 5 of 5 each; they took about as long and DeepSeek V4 Pro cost 12.9× less.

  • GPT-6 Astra vs DeepSeek V4.1 Flash: 5 of 5 each; DeepSeek V4.1 Flash answered 1.5× sooner at the median and cost 36.8× less.

  • GPT-6 Sol vs DeepSeek V4 Pro: 5 of 5 each; GPT-6 Sol answered 1.3× sooner at the median and DeepSeek V4 Pro cost 2.7× less. GPT-6 Sol vs DeepSeek V4 Pro, on every job.

  • GPT-6 Sol vs DeepSeek V4.1 Flash: 5 of 5 each; DeepSeek V4.1 Flash answered 1.3× sooner at the median and cost 7.8× less.

  • GPT-6 Luna vs DeepSeek V4 Pro: GPT-6 Luna 4 of 5, DeepSeek V4 Pro 5; GPT-6 Luna answered 1.2× sooner at the median and cost 6.7× less. GPT-6 Luna vs DeepSeek V4 Pro, on every job.

  • GPT-6 Luna vs DeepSeek V4.1 Flash: GPT-6 Luna 4 of 5, DeepSeek V4.1 Flash 5; DeepSeek V4.1 Flash answered 1.3× sooner at the median and GPT-6 Luna cost 2.4× less. GPT-6 Luna vs DeepSeek V4.1 Flash, on every job.

How the math replies are scored

Every GPT and DeepSeek reply above was checked the same way as every other model's, by the rules published with the prompts: how the math prompts are scored, and each one in full.

The differences at a glance

What follows from each model's facts in our catalog.

  • The lineups

    GPT: 3 models, GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. DeepSeek: 2 models, DeepSeek V4 Pro and DeepSeek V4.1 Flash.

  • Price per message

    The least expensive GPT model is GPT-6 Luna (60 messages a day on Pro); the least expensive DeepSeek model is DeepSeek V4.1 Flash (60 messages a day on Pro).

  • Context window

    GPT goes up to 1.05M tokens (GPT-6 Astra); DeepSeek up to 1.05M tokens (DeepSeek V4 Pro).

  • Images and PDFs

    DeepSeek V4 Pro doesn't read images. DeepSeek V4 Pro and DeepSeek V4.1 Flash get a PDF's text rather than the file itself.

  • On the Free plan

    Free's one-time trial of 5 messages covers GPT-6 Sol, GPT-6 Luna, DeepSeek V4 Pro, and DeepSeek V4.1 Flash. Paid plans have every model, with messages every month.

Every GPT and DeepSeek model's context window, files and API price: DeepSeek vs ChatGPT.

Where your messages go

In llmwise, a message to GPT goes to its maker, OpenAI, or through OpenRouter when llmwise can't reach the maker directly. DeepSeek models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.

Questions

Which is better, ChatGPT or DeepSeek for math?

In our math test runs on September 27, 2026, GPT's 3 models passed 14 of 15; GPT-6 Astra and GPT-6 Sol each passed 5 of 5. DeepSeek's 2 models passed 10 of 10; DeepSeek V4 Pro and DeepSeek V4.1 Flash each passed 5 of 5. GPT-6 Sol and DeepSeek V4.1 Flash each passed 5 of the 5 prompts, so these math prompts don't split GPT and DeepSeek; the replies on this page show how they differ.

Is GPT or DeepSeek cheaper?

In llmwise, the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro), and the least expensive DeepSeek model is DeepSeek V4.1 Flash (60 messages a day on Pro). At API list prices (September 2026), a typical message of 4,000 tokens in and 700 out costs $0.0008 on GPT-6 Luna and $0.0011 on DeepSeek V4.1 Flash.

Can I use GPT and DeepSeek in the same chat?

Yes. Ask GPT-6 Sol a question, then switch the picker to DeepSeek V4.1 Flash and ask again: DeepSeek V4.1 Flash sees the whole conversation, GPT-6 Sol's answer included.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.