Skip to content

Comparison

Claude Haiku vs ChatGPT

In our test runs on September 29, 2026, the same 50 prompts across 10 jobs: Claude Haiku's one model passed 43 of 50 replies and GPT's 4 models passed 186 of 200 replies. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself.

Based on 250 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

Job by job, both families' best models shared the top result on 6 of the 10 jobs; GPT's alone had it on coding, writing, data analysis, and customer support.

Claude Haiku vs GPT, job by job

On each job, Claude Haiku's pick against GPT's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.

  • Coding: Claude Haiku 4.5 passed 3 of 5 and GPT-6 Luna 5 of 5. Claude Haiku 4.5 answered 1.9× sooner at the median, 2.3 s against 4.6 s. GPT-6 Luna cost 9.6× less, $0.0137 against $0.0014 for the 5 replies. Claude Haiku 4.5's replies ran 84% longer, in tokens of reply, thinking not counted.

  • Data analysis: Claude Haiku 4.5 passed 3 of 5 and GPT-6 Luna 5 of 5. Their median waits were close, 2.7 s against 2.5 s. GPT-6 Luna cost 14.2× less, $0.0125 against $0.0009 for the 5 replies. Claude Haiku 4.5's replies ran 403% longer, in tokens of reply, thinking not counted.

  • Writing: Claude Haiku 4.5 passed 4 of 5 and GPT-6 Luna 5 of 5. Their median waits were close, 2.2 s against 2.4 s. GPT-6 Luna cost 11.4× less, $0.0057 against $0.0005 for the 5 replies. Claude Haiku 4.5's replies ran 31% longer, in tokens of reply, thinking not counted.

  • Customer support: Claude Haiku 4.5 passed 3 of 5 and GPT-6 Luna 4 of 5. GPT-6 Luna answered 1.9× sooner at the median, 2.7 s against 1.4 s. GPT-6 Luna cost 17.0× less, $0.0073 against $0.0004 for the 5 replies. Claude Haiku 4.5's replies ran 160% longer, in tokens of reply, thinking not counted.

  • Math: Claude Haiku 4.5 passed 5 of 5 and GPT-6.1 Sol 5 of 5. GPT-6.1 Sol answered 1.3× sooner at the median, 2.2 s against 1.8 s. GPT-6.1 Sol cost 1.8× less, $0.0077 against $0.0043 for the 5 replies. Claude Haiku 4.5's replies ran 188% longer, in tokens of reply, thinking not counted.

  • Summarization: Claude Haiku 4.5 passed 5 of 5 and GPT-6.1 Sol 5 of 5. Claude Haiku 4.5 answered 1.2× sooner at the median, 1.7 s against 2.0 s. They cost about the same, $0.0051 against $0.0050 for the 5 replies. Their replies ran to about the same length.

  • Translation: Claude Haiku 4.5 passed 5 of 5 and GPT-6 Luna 5 of 5. GPT-6 Luna answered 1.2× sooner at the median, 1.8 s against 1.6 s. GPT-6 Luna cost 15.9× less, $0.0074 against $0.0005 for the 5 replies. Claude Haiku 4.5's replies ran 105% longer, in tokens of reply, thinking not counted.

  • SQL: Claude Haiku 4.5 passed 5 of 5 and GPT-6 Luna 5 of 5. Their median waits were close, 1.2 s against 1.2 s. GPT-6 Luna cost 11.1× less, $0.0049 against $0.0004 for the 5 replies. Claude Haiku 4.5's replies ran 13% longer, in tokens of reply, thinking not counted.

  • RAG and answering from documents: Claude Haiku 4.5 passed 5 of 5 and GPT-6 Luna 5 of 5. Their median waits were close, 1.2 s against 1.3 s. GPT-6 Luna cost 13.1× less, $0.0050 against $0.0004 for the 5 replies. Claude Haiku 4.5's replies ran 133% longer, in tokens of reply, thinking not counted.

  • Agents and tool use: Claude Haiku 4.5 passed 5 of 5 and GPT-6 Luna 5 of 5. Claude Haiku 4.5 answered 1.7× sooner at the median, 1.1 s against 1.9 s. GPT-6 Luna cost 11.7× less, $0.0047 against $0.0004 for the 5 replies. Claude Haiku 4.5's replies ran 101% longer, in tokens of reply, thinking not counted.

The 6 prompts only one of Claude Haiku and GPT passed

Where one family's pick passed a prompt and the other's didn't, in each check's own words.

  • Evaluate an arithmetic expression, no eval (coding): GPT-6 Luna passed and Claude Haiku 4.5 didn't. Claude Haiku 4.5: 14 of 15 tests passed. First failure: evaluate("-2 ^ 2"): Expected values to be strictly deep-equal: GPT-6 Luna: All 15 tests passed.

  • Parse CSV with quoted fields (coding): GPT-6 Luna passed and Claude Haiku 4.5 didn't. Claude Haiku 4.5: 5 of 8 tests passed. First failure: parseCsv('"line1\nline2",end\r\nnext,row\r\n'): Expected values to be strictly deep-equal: GPT-6 Luna: All 8 tests passed.

  • Average order value in August (data analysis): GPT-6 Luna passed and Claude Haiku 4.5 didn't. Claude Haiku 4.5: Final answer 275.00; expected 300. GPT-6 Luna: Final answer 300.00: right.

  • Correlation between ad spend and sign-ups (data analysis): GPT-6 Luna passed and Claude Haiku 4.5 didn't. Claude Haiku 4.5: Final answer 0.99; expected 0.97. GPT-6 Luna: Final answer 0.97: right.

  • A product announcement with five rules (writing): GPT-6 Luna passed and Claude Haiku 4.5 didn't. Claude Haiku 4.5: Graded 3.7 of 5 on average (lowest 3). GPT-6 Luna: Graded 4.0 of 5 on average (lowest 3).

  • A message with a planted instruction (customer support): GPT-6 Luna passed and Claude Haiku 4.5 didn't. Claude Haiku 4.5: Graded 4.3 of 5 on average (lowest 3); but 165 words, over the 150 allowed. GPT-6 Luna: Graded 4.7 of 5 on average (lowest 4).

Each Claude Haiku model against each GPT model

Every Claude Haiku model against every GPT model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.

  • Claude Haiku 4.5 vs GPT-6 Astra: 43 and 48 of 50; 7 prompts split them; Claude Haiku 4.5's replies cost 7.9× less in all, and Claude Haiku 4.5 answered sooner on 41, GPT-6 Astra on 4.

  • Claude Haiku 4.5 vs GPT-6.1 Sol: 43 and 46 of 50; 9 prompts split them; GPT-6.1 Sol's replies cost 1.2× less in all, and Claude Haiku 4.5 answered sooner on 30, GPT-6.1 Sol on 10.

  • Claude Haiku 4.5 vs GPT-6 Sol: 43 and 45 of 50; 10 prompts split them; Claude Haiku 4.5's replies cost 1.7× less in all, and Claude Haiku 4.5 answered sooner on 30, GPT-6 Sol on 12. Claude Haiku 4.5 vs GPT-6 Sol.

  • Claude Haiku 4.5 vs GPT-6 Luna: 43 and 47 of 50; 8 prompts split them; GPT-6 Luna's replies cost 12.6× less in all, and GPT-6 Luna answered sooner on 21, Claude Haiku 4.5 on 20. Claude Haiku 4.5 vs GPT-6 Luna.

One coding prompt, both replies

Claude Haiku 4.5 and GPT-6 Luna on a coding prompt, the job where they differed most: both replies as they came.

A hard prompt: “Evaluate an arithmetic expression, no eval”, in full.

  • Claude Haiku 4.5

    Failed: 14 of 15 tests passed. First failure: evaluate("-2 ^ 2"): Expected values to be strictly deep-equal:

    ```javascript
    function evaluate(expression) {
      const tokens = tokenize(expression);
      const parser = new Parser(tokens);
      const result = parser.parseExpression();
      
      if (parser.pos !== tokens.length) {
        throw new Error("Unexpected tokens after expression");
      }
      
      return result;
    }
    
    function tokenize(expression) {
      const tokens = [];
      let i = 0;
      
      while (i < expression.length) {
        const ch = expression[i];
        
        if (/\s/.test(ch)) {
          i++;
          continue;
        }
        
        if (/\d/.test(ch)) {
          let num = '';
          while (i < expression.length && /[\d.]/.test(expression[i])) {
            num += expression[i];
            i++;
          }
          const parsed = parseFloat(num);
          if (isNaN(parsed)) {
            throw new Error(`Invalid number: ${num}`);
          }
          tokens.push({ type: 'NUMBER', value: parsed });
          continue;
        }
        
        if ('+-*/^()'.includes(ch)) {
          tokens.push({ type: ch, value: ch });
          i++;
          continue;
        }
        
        throw new Error(`Unexpected character: ${ch}`);
      }
      
      return tokens;
    }
    
    class Parser {
      constructor(tokens) {
        this.tokens = tokens;
        this.pos = 0;
      }
      
      peek() {
        return this.tokens[this.pos];
      }
      
      consume() {
        return this.tokens[this.pos++];
      }
      
      parseExpression() {
        return this.parseAddSub();
      }
      
      parseAddSub() {
        let left = this.parseMulDiv();
        
        while (this.peek() && (this.peek().type === '+' || this.peek().type === '-')) {
          const op = this.consume().type;
          const right = this.parseMulDiv();
          left = op === '+' ? left + right : left - right;
        }
        
        return left;
      }
      
      parseMulDiv() {
        let left = this.parsePower();
        
        while (this.peek() && (this.peek().type === '*' || this.peek().type === '/')) {
          const op = this.consume().type;
          const right = this.parsePower();
          left = op === '*' ? left * right : left / right;
        }
        
        return left;
      }
      
      parsePower() {
        let left = this.parseUnary();
        
        if (this.peek() && this.peek().type === '^') {
          this.consume();
          const right = this.parsePower();
          left = Math.pow(left, right);
        }
        
        return left;
      }
      
      parseUnary() {
        if (this.peek() && this.peek().type === '-') {
          this.consume();
          return -this.parseUnary();
        }
        
        if (this.peek() && this.peek().type === '+') {
          this.consume();
          return this.parseUnary();
        }
        
        return this.parsePrimary();
      }
      
      parsePrimary() {
        const token = this.peek();
        
        if (!token) {
          throw new Error("Unexpected end of expression");
        }
        
        if (token.type === 'NUMBER') {
          this.consume();
          return token.value;
        }
        
        if (token.type === '(') {
          this.consume();
          const result = this.parseExpression();
          
          if (!this.peek() || this.peek().type !== ')') {
            throw new Error("Missing closing parenthesis");
          }
          
          this.consume();
          return result;
        }
        
        throw new Error(`Unexpected token: ${token.value}`);
      }
    }
    ```

    599 tokens in, 1,007 out · 7.6 s · $0.0056 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

  • GPT-6 Luna

    Passed: All 15 tests passed.

    ```javascript
    function evaluate(expression) {
      if (typeof expression !== "string") {
        throw new Error("Expression must be a string");
      }
    
      let index = 0;
    
      function skipSpaces() {
        while (index < expression.length && /\s/.test(expression[index])) {
          index++;
        }
      }
    
      function parsePrimary() {
        skipSpaces();
    
        if (expression[index] === "(") {
          index++;
          const value = parseExpression();
          skipSpaces();
          if (expression[index] !== ")") throw new Error("Missing closing parenthesis");
          index++;
          return value;
        }
    
        const start = index;
        let digitsBefore = 0;
        while (/\d/.test(expression[index] || "")) {
          index++;
          digitsBefore++;
        }
    
        let digitsAfter = 0;
        if (expression[index] === ".") {
          index++;
          while (/\d/.test(expression[index] || "")) {
            index++;
            digitsAfter++;
          }
        }
    
        if (digitsBefore === 0 && digitsAfter === 0) {
          throw new Error("Expected a number or parenthesized expression");
        }
    
        return Number(expression.slice(start, index));
      }
    
      // Unary minus has lower precedence than exponentiation, but is allowed
      // on an exponent's right-hand side.
      function parseUnary() {
        skipSpaces();
        if (expression[index] === "-") {
          index++;
          return -parseUnary();
        }
        return parsePower();
      }
    
      function parsePower() {
        let value = parsePrimary();
        skipSpaces();
    
        if (expression[index] === "^") {
          index++;
          value **= parseUnary();
        }
    
        return value;
      }
    
      function parseProduct() {
        let value = parseUnary();
    
        while (true) {
          skipSpaces();
          const operator = expression[index];
          if (operator !== "*" && operator !== "/") break;
    
          index++;
          const right = parseUnary();
          value = operator === "*" ? value * right : value / right;
        }
    
        return value;
      }
    
      function parseExpression() {
        let value = parseProduct();
    
        while (true) {
          skipSpaces();
          const operator = expression[index];
          if (operator !== "+" && operator !== "-") break;
    
          index++;
          const right = parseProduct();
          value = operator === "+" ? value + right : value - right;
        }
    
        return value;
      }
    
      const result = parseExpression();
      skipSpaces();
      if (index !== expression.length) throw new Error("Unexpected input");
    
      return result;
    }
    ```

    549 tokens in, 1,007 out (455 of them reasoning) · 8.8 s · $0.0006 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

Every model, every job

All 5 Claude Haiku and GPT models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.

Every Claude Haiku and GPT model across our test runs
ModelPassedHard onesCost per replyOn Pro
Claude Haiku 4.5Anthropic43 of 5015 of 20$0.0015Up to 250 a month
GPT-6 AstraOpenAI48 of 5020 of 20$0.0117Up to 31 a month
GPT-6.1 SolOpenAI46 of 5019 of 20$0.0012Up to 125 a month
GPT-6 SolOpenAI45 of 5019 of 20$0.0026Up to 125 a month
GPT-6 LunaOpenAI47 of 5020 of 20$0.00012Up to 60 a day

Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Their own subscriptions

Each company's own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.

llmwise Pro, $20 a month, has all 5 of these models in one chat, on one monthly allowance. On it: Claude Haiku 4.5 up to 250 messages a month and GPT-6.1 Sol up to 125.

The lineups at a glance

What follows from each model's facts in our catalog.

  • The lineups

    Claude Haiku: one model, Claude Haiku 4.5. GPT: 4 models, GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna.

  • Price per message

    Claude Haiku's one model is Claude Haiku 4.5 (250 messages a month on Pro); the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro).

  • Context window

    Claude Haiku goes up to 200K tokens (Claude Haiku 4.5); GPT up to 1.05M tokens (GPT-6 Astra).

  • Images and PDFs

    Every model here reads images. Every model here takes a PDF as a whole file.

  • On the Free plan

    Free's one-time trial of 5 messages covers Claude Haiku 4.5, GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna. Paid plans have every model, with messages every month.

Model by model

Two named models side by side, prompt by prompt, each with its messages on every plan.

More head-to-heads

Each of Claude Haiku and GPT against the other families, every job from the same test runs.

Where your messages go

In llmwise, a message to Claude or GPT goes to the model's maker, or through OpenRouter when llmwise can't reach the maker directly. The Privacy Policy has the details.

Bar chart: Prompts passed in our test runs, all 10 jobs. GPT-6 Astra: 48 of 50; GPT-6 Luna: 47 of 50; GPT-6.1 Sol: 46 of 50; GPT-6 Sol: 45 of 50; Claude Haiku 4.5: 43 of 50.
Our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026: the same prompts for every model, each reply checked the same way.

Questions

Which is better, Claude Haiku or ChatGPT?

In our test runs on September 29, 2026, the same 50 prompts across 10 jobs: Claude Haiku's one model passed 43 of 50 replies and GPT's 4 models passed 186 of 200 replies. Job by job, both families' best models shared the top result on 6 of the 10 jobs; GPT's alone had it on coding, writing, data analysis, and customer support.

Which is cheaper, Claude Haiku or GPT?

In llmwise, the least expensive Claude Haiku model is Claude Haiku 4.5 (250 messages a month on Pro), and the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro). At API list prices (October 2026), a typical message of 4,000 tokens in and 700 out costs $0.0075 on Claude Haiku 4.5 and $0.0008 on GPT-6 Luna.

Can I use Claude Haiku and GPT in the same chat?

Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.