Skip to content

Comparison

DeepSeek vs ChatGPT

In our test runs on October 8, 2026, the same 50 prompts across 10 jobs: DeepSeek's 2 models passed 97 of 100 replies and GPT's 4 models passed 186 of 200 replies. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself.

Based on 300 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

Job by job, both families' best models shared the top result on 10 of the 10 jobs.

DeepSeek vs GPT, job by job

On each job, DeepSeek's pick against GPT's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.

  • Coding: DeepSeek V4.1 Flash passed 5 of 5 and GPT-6 Luna 5 of 5. DeepSeek V4.1 Flash answered 1.8× sooner at the median, 2.5 s against 4.6 s. GPT-6 Luna cost 6.4× less, $0.0092 against $0.0014 for the 5 replies. DeepSeek V4.1 Flash's replies ran 62% longer, in tokens of reply, thinking not counted.

  • Writing: DeepSeek V4 Pro passed 5 of 5 and GPT-6 Luna 5 of 5. GPT-6 Luna answered 1.6× sooner at the median, 3.9 s against 2.4 s. GPT-6 Luna cost 10.8× less, $0.0054 against $0.0005 for the 5 replies. Their replies ran to about the same length.

  • Math: DeepSeek V4.1 Flash passed 5 of 5 and GPT-6.1 Sol 5 of 5. DeepSeek V4.1 Flash answered 1.9× sooner at the median, 0.9 s against 1.8 s. DeepSeek V4.1 Flash cost 3.4× less, $0.0012 against $0.0043 for the 5 replies. DeepSeek V4.1 Flash's replies ran 15% longer, in tokens of reply, thinking not counted.

  • Summarization: DeepSeek V4.1 Flash passed 5 of 5 and GPT-6.1 Sol 5 of 5. DeepSeek V4.1 Flash answered 3.9× sooner at the median, 0.5 s against 2.0 s. DeepSeek V4.1 Flash cost 5.0× less, $0.0010 against $0.0050 for the 5 replies. GPT-6.1 Sol's replies ran 11% longer, in tokens of reply, thinking not counted.

  • Data analysis: DeepSeek V4.1 Flash passed 5 of 5 and GPT-6 Luna 5 of 5. DeepSeek V4.1 Flash answered 2.9× sooner at the median, 0.9 s against 2.5 s. GPT-6 Luna cost 4.8× less, $0.0042 against $0.0009 for the 5 replies. DeepSeek V4.1 Flash's replies ran 111% longer, in tokens of reply, thinking not counted.

  • Customer support: DeepSeek V4 Pro passed 4 of 5 and GPT-6 Luna 4 of 5. GPT-6 Luna answered 2.9× sooner at the median, 4.1 s against 1.4 s. GPT-6 Luna cost 17.8× less, $0.0077 against $0.0004 for the 5 replies. DeepSeek V4 Pro's replies ran 61% longer, in tokens of reply, thinking not counted.

  • Translation: DeepSeek V4.1 Flash passed 5 of 5 and GPT-6 Luna 5 of 5. DeepSeek V4.1 Flash answered 1.2× sooner at the median, 1.3 s against 1.6 s. GPT-6 Luna cost 6.3× less, $0.0029 against $0.0005 for the 5 replies. DeepSeek V4.1 Flash's replies ran 63% longer, in tokens of reply, thinking not counted.

  • SQL: DeepSeek V4.1 Flash passed 5 of 5 and GPT-6 Luna 5 of 5. DeepSeek V4.1 Flash answered 2.9× sooner at the median, 0.4 s against 1.2 s. GPT-6 Luna cost 2.3× less, $0.0010 against $0.0004 for the 5 replies. Their replies ran to about the same length.

  • RAG and answering from documents: DeepSeek V4.1 Flash passed 5 of 5 and GPT-6 Luna 5 of 5. DeepSeek V4.1 Flash answered 3.7× sooner at the median, 0.3 s against 1.3 s. GPT-6 Luna cost 2.2× less, $0.0008 against $0.0004 for the 5 replies. DeepSeek V4.1 Flash's replies ran 120% longer, in tokens of reply, thinking not counted.

  • Agents and tool use: DeepSeek V4.1 Flash passed 5 of 5 and GPT-6 Luna 5 of 5. DeepSeek V4.1 Flash answered 5.4× sooner at the median, 0.4 s against 1.9 s. GPT-6 Luna cost 2.5× less, $0.0010 against $0.0004 for the 5 replies. DeepSeek V4.1 Flash's replies ran 21% longer, in tokens of reply, thinking not counted.

On every prompt, DeepSeek's pick and GPT's both passed or both failed: these prompts don't split them.

Each DeepSeek model against each GPT model

Every DeepSeek model against every GPT model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.

  • DeepSeek V4 Pro vs GPT-6 Astra: 49 and 48 of 50; one prompt split them; DeepSeek V4 Pro's replies cost 3.2× less in all, and GPT-6 Astra answered sooner on 23, DeepSeek V4 Pro on 18.

  • DeepSeek V4 Pro vs GPT-6.1 Sol: 49 and 46 of 50; 3 prompts split them; GPT-6.1 Sol's replies cost 3.1× less in all, and GPT-6.1 Sol answered sooner on 32, DeepSeek V4 Pro on 11.

  • DeepSeek V4 Pro vs GPT-6 Sol: 49 and 45 of 50; 4 prompts split them; GPT-6 Sol's replies cost 1.4× less in all, and GPT-6 Sol answered sooner on 32, DeepSeek V4 Pro on 12. GPT-6 Sol vs DeepSeek V4 Pro.

  • DeepSeek V4 Pro vs GPT-6 Luna: 49 and 47 of 50; 2 prompts split them; GPT-6 Luna's replies cost 31.0× less in all, and GPT-6 Luna answered sooner on 39, DeepSeek V4 Pro on 8. GPT-6 Luna vs DeepSeek V4 Pro.

  • DeepSeek V4.1 Flash vs GPT-6 Astra: 48 and 48 of 50; 4 prompts split them; DeepSeek V4.1 Flash's replies cost 22.5× less in all, and DeepSeek V4.1 Flash answered sooner on 49, GPT-6 Astra on 0.

  • DeepSeek V4.1 Flash vs GPT-6.1 Sol: 48 and 46 of 50; 4 prompts split them; DeepSeek V4.1 Flash's replies cost 2.3× less in all, and DeepSeek V4.1 Flash answered sooner on 45, GPT-6.1 Sol on 2.

  • DeepSeek V4.1 Flash vs GPT-6 Sol: 48 and 45 of 50; 5 prompts split them; DeepSeek V4.1 Flash's replies cost 5.0× less in all, and DeepSeek V4.1 Flash answered sooner on 45, GPT-6 Sol on 3.

  • DeepSeek V4.1 Flash vs GPT-6 Luna: 48 and 47 of 50; 5 prompts split them; GPT-6 Luna's replies cost 4.4× less in all, and DeepSeek V4.1 Flash answered sooner on 41, GPT-6 Luna on 6. GPT-6 Luna vs DeepSeek V4.1 Flash.

One coding prompt, both replies

DeepSeek V4.1 Flash and GPT-6 Luna on a coding prompt: both replies as they came.

A hard prompt: “Evaluate an arithmetic expression, no eval”, in full.

  • DeepSeek V4.1 Flash

    Passed: All 15 tests passed.

    ```javascript
    function evaluate(expression) {
      if (typeof expression !== 'string') {
        throw new Error('Expression must be a string');
      }
    
      // ---- Tokenizer ----
      const tokens = [];
      let i = 0;
      const isDigit = (c) => c >= '0' && c <= '9';
    
      while (i < expression.length) {
        const ch = expression[i];
    
        if (ch === ' ' || ch === '\t') {
          i++;
          continue;
        }
    
        // number: digits with an optional single decimal point (at least one digit)
        if (isDigit(ch) || ch === '.') {
          let j = i;
          let digits = 0;
          while (j < expression.length && isDigit(expression[j])) {
            j++;
            digits++;
          }
          if (j < expression.length && expression[j] === '.') {
            j++;
            while (j < expression.length && isDigit(expression[j])) {
              j++;
              digits++;
            }
          }
          if (digits === 0) {
            throw new Error('Invalid number in expression');
          }
          tokens.push({ type: 'num', value: parseFloat(expression.slice(i, j)) });
          i = j;
          continue;
        }
    
        if ('+-*/^()'.indexOf(ch) !== -1) {
          tokens.push({ type: ch });
          i++;
          continue;
        }
    
        throw new Error('Unexpected character: ' + ch);
      }
    
      if (tokens.length === 0) {
        throw new Error('Empty expression');
      }
    
      // ---- Parser (recursive descent) ----
      let pos = 0;
      const peek = () => tokens[pos];
      const isOp = (tok, ch) => tok !== undefined && tok.type === ch;
    
      // expr := term (('+' | '-') term)*          left-associative
      function parseExpression() {
        let value = parseTerm();
        while (isOp(peek(), '+') || isOp(peek(), '-')) {
          const op = tokens[pos++].type;
          const rhs = parseTerm();
          value = op === '+' ? value + rhs : value - rhs;
        }
        return value;
      }
    
      // term := unary (('*' | '/') unary)*        left-associative
      function parseTerm() {
        let value = parseUnary();
        while (isOp(peek(), '*') || isOp(peek(), '/')) {
          const op = tokens[pos++].type;
          const rhs = parseUnary();
          value = op === '*' ? value * rhs : value / rhs;
        }
        return value;
      }
    
      // unary := ('-' | '+') unary | power
      // unary minus is looser than '^', so -2^2 === -(2^2) === -4
      function parseUnary() {
        if (isOp(peek(), '-')) {
          pos++;
          return -parseUnary();
        }
        if (isOp(peek(), '+')) {
          pos++;
          return parseUnary();
        }
        return parsePower();
      }
    
      // power := primary ('^' unary)?             right-associative
      // the exponent is a unary, which permits 2^-1 and 2^3^2 = 2^(3^2)
      function parsePower() {
        const base = parsePrimary();
        if (isOp(peek(), '^')) {
          pos++;
          const exponent = parseUnary();
          return Math.pow(base, exponent);
        }
        return base;
      }
    
      // primary := number | '(' expr ')'
      function parsePrimary() {
        const tok = peek();
        if (tok === undefined) {
          throw new Error('Unexpected end of expression');
        }
        if (tok.type === 'num') {
          pos++;
          return tok.value;
        }
        if (tok.type === '(') {
          pos++;
          const value = parseExpression();
          if (!isOp(peek(), ')')) {
            throw new Error('Missing closing parenthesis');
          }
          pos++;
          return value;
        }
        throw new Error('Unexpected token: ' + tok.type);
      }
    
      const result = parseExpression();
      if (pos !== tokens.length) {
        throw new Error('Unexpected token: ' + tokens[pos].type);
      }
      return result;
    }
    ```

    591 tokens in, 2,442 out (1,491 of them reasoning) · 5.8 s · $0.0030 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Together ·

  • GPT-6 Luna

    Passed: All 15 tests passed.

    ```javascript
    function evaluate(expression) {
      if (typeof expression !== "string") {
        throw new Error("Expression must be a string");
      }
    
      let index = 0;
    
      function skipSpaces() {
        while (index < expression.length && /\s/.test(expression[index])) {
          index++;
        }
      }
    
      function parsePrimary() {
        skipSpaces();
    
        if (expression[index] === "(") {
          index++;
          const value = parseExpression();
          skipSpaces();
          if (expression[index] !== ")") throw new Error("Missing closing parenthesis");
          index++;
          return value;
        }
    
        const start = index;
        let digitsBefore = 0;
        while (/\d/.test(expression[index] || "")) {
          index++;
          digitsBefore++;
        }
    
        let digitsAfter = 0;
        if (expression[index] === ".") {
          index++;
          while (/\d/.test(expression[index] || "")) {
            index++;
            digitsAfter++;
          }
        }
    
        if (digitsBefore === 0 && digitsAfter === 0) {
          throw new Error("Expected a number or parenthesized expression");
        }
    
        return Number(expression.slice(start, index));
      }
    
      // Unary minus has lower precedence than exponentiation, but is allowed
      // on an exponent's right-hand side.
      function parseUnary() {
        skipSpaces();
        if (expression[index] === "-") {
          index++;
          return -parseUnary();
        }
        return parsePower();
      }
    
      function parsePower() {
        let value = parsePrimary();
        skipSpaces();
    
        if (expression[index] === "^") {
          index++;
          value **= parseUnary();
        }
    
        return value;
      }
    
      function parseProduct() {
        let value = parseUnary();
    
        while (true) {
          skipSpaces();
          const operator = expression[index];
          if (operator !== "*" && operator !== "/") break;
    
          index++;
          const right = parseUnary();
          value = operator === "*" ? value * right : value / right;
        }
    
        return value;
      }
    
      function parseExpression() {
        let value = parseProduct();
    
        while (true) {
          skipSpaces();
          const operator = expression[index];
          if (operator !== "+" && operator !== "-") break;
    
          index++;
          const right = parseProduct();
          value = operator === "+" ? value + right : value - right;
        }
    
        return value;
      }
    
      const result = parseExpression();
      skipSpaces();
      if (index !== expression.length) throw new Error("Unexpected input");
    
      return result;
    }
    ```

    549 tokens in, 1,007 out (455 of them reasoning) · 8.8 s · $0.0006 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

Every model, every job

All 6 DeepSeek and GPT models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.

Every DeepSeek and GPT model across our test runs
ModelPassedHard onesCost per replyOn Pro
DeepSeek V4 ProDeepSeek49 of 5020 of 20$0.0037Up to 250 a month
DeepSeek V4.1 FlashDeepSeek48 of 5019 of 20$0.00052Up to 60 a day
GPT-6 AstraOpenAI48 of 5020 of 20$0.0117Up to 31 a month
GPT-6.1 SolOpenAI46 of 5019 of 20$0.0012Up to 125 a month
GPT-6 SolOpenAI45 of 5019 of 20$0.0026Up to 125 a month
GPT-6 LunaOpenAI47 of 5020 of 20$0.00012Up to 60 a day

Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Its own subscription

The one company here with its own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.

llmwise Pro, $20 a month, has all 6 of these models in one chat, on one monthly allowance. On it: DeepSeek V4 Pro up to 250 messages a month and GPT-6.1 Sol up to 125.

The lineups at a glance

What follows from each model's facts in our catalog.

  • The lineups

    DeepSeek: 2 models, DeepSeek V4 Pro and DeepSeek V4.1 Flash. GPT: 4 models, GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna.

  • Price per message

    The least expensive DeepSeek model is DeepSeek V4.1 Flash (60 messages a day on Pro); the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro).

  • Context window

    DeepSeek goes up to 1.05M tokens (DeepSeek V4 Pro); GPT up to 1.05M tokens (GPT-6 Astra).

  • Images and PDFs

    DeepSeek V4 Pro doesn't read images. DeepSeek V4 Pro and DeepSeek V4.1 Flash get a PDF's text rather than the file itself.

  • On the Free plan

    Free's one-time trial of 5 messages covers DeepSeek V4 Pro, DeepSeek V4.1 Flash, GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna. Paid plans have every model, with messages every month.

Model by model

Two named models side by side, prompt by prompt, each with its messages on every plan.

More head-to-heads

Each of DeepSeek and GPT against the other families, every job from the same test runs.

Where your messages go

In llmwise, a message to GPT goes to its maker, OpenAI, or through OpenRouter when llmwise can't reach the maker directly. DeepSeek models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.

Bar chart: Prompts passed in our test runs, all 10 jobs. DeepSeek V4 Pro: 49 of 50; GPT-6 Astra: 48 of 50; DeepSeek V4.1 Flash: 48 of 50; GPT-6 Luna: 47 of 50; GPT-6.1 Sol: 46 of 50; GPT-6 Sol: 45 of 50.
Our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026: the same prompts for every model, each reply checked the same way.

Questions

Which is better, DeepSeek or ChatGPT?

In our test runs on October 8, 2026, the same 50 prompts across 10 jobs: DeepSeek's 2 models passed 97 of 100 replies and GPT's 4 models passed 186 of 200 replies. Job by job, both families' best models shared the top result on 10 of the 10 jobs.

Which is cheaper, DeepSeek or GPT?

In llmwise, the least expensive DeepSeek model is DeepSeek V4.1 Flash (60 messages a day on Pro), and the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro). At API list prices (October 2026), a typical message of 4,000 tokens in and 700 out costs $0.0020 on DeepSeek V4.1 Flash and $0.0008 on GPT-6 Luna.

Can I use DeepSeek and GPT in the same chat?

Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.