Skip to content

Comparison

Claude vs Gemini

In our test runs on October 7, 2026, the same 50 prompts across 10 jobs: Claude's 6 models passed 275 of 300 replies and Gemini's 2 models passed 95 of 100 replies. On llmwise Pro, Claude Sonnet 5.5 up to 125 messages a month and Gemini 3.1 Pro (preview) up to 125.

Based on 400 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

Job by job, both families' best models shared the top result on 10 of the 10 jobs.

Claude vs Gemini, job by job

On each job, Claude's pick against Gemini's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.

  • Coding: Claude Haiku 5.5 passed 5 of 5 and Gemini 3.8 Flash 5 of 5. Their median waits were close, 3.9 s against 4.2 s. Claude Haiku 5.5 cost 4.3× less, $0.0023 against $0.0098 for the 5 replies. Their replies ran to about the same length.

  • Writing: Claude Sonnet 5 passed 4 of 5 and Gemini 3.1 Pro (preview) 4 of 5. Claude Sonnet 5 answered 2.0× sooner at the median, 4.5 s against 9.1 s. Claude Sonnet 5 cost 2.7× less, $0.0171 against $0.0459 for the 5 replies. Claude Sonnet 5's replies ran 75% longer, in tokens of reply, thinking not counted.

  • Math: Claude Haiku 5.5 passed 5 of 5 and Gemini 3.8 Flash 5 of 5. Claude Haiku 5.5 answered 2.7× sooner at the median, 1.7 s against 4.7 s. Claude Haiku 5.5 cost 9.7× less, $0.0007 against $0.0071 for the 5 replies. Claude Haiku 5.5's replies ran 14% longer, in tokens of reply, thinking not counted.

  • Summarization: Claude Haiku 4.5 passed 5 of 5 and Gemini 3.8 Flash 5 of 5. Claude Haiku 4.5 answered 2.2× sooner at the median, 1.7 s against 3.7 s. Gemini 3.8 Flash cost 1.3× less, $0.0051 against $0.0038 for the 5 replies. Their replies ran to about the same length.

  • Data analysis: Claude Haiku 5.5 passed 5 of 5 and Gemini 3.1 Pro (preview) 5 of 5. Claude Haiku 5.5 answered 2.6× sooner at the median, 3.3 s against 8.6 s. Claude Haiku 5.5 cost 43.1× less, $0.0018 against $0.0773 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 73% longer, in tokens of reply, thinking not counted.

  • Customer support: Claude Sonnet 5 passed 5 of 5 and Gemini 3.1 Pro (preview) 5 of 5. Claude Sonnet 5 answered 2.5× sooner at the median, 3.5 s against 8.7 s. Claude Sonnet 5 cost 2.9× less, $0.0182 against $0.0529 for the 5 replies. Claude Sonnet 5's replies ran 61% longer, in tokens of reply, thinking not counted.

  • Translation: Claude Haiku 5.5 passed 5 of 5 and Gemini 3.8 Flash 5 of 5. Claude Haiku 5.5 answered 2.1× sooner at the median, 1.4 s against 3.0 s. Claude Haiku 5.5 cost 4.2× less, $0.0009 against $0.0037 for the 5 replies. Claude Haiku 5.5's replies ran 100% longer, in tokens of reply, thinking not counted.

  • SQL: Claude Haiku 5.5 passed 5 of 5 and Gemini 3.8 Flash 5 of 5. Claude Haiku 5.5 answered 2.9× sooner at the median, 1.3 s against 3.7 s. Claude Haiku 5.5 cost 3.4× less, $0.0011 against $0.0037 for the 5 replies. Claude Haiku 5.5's replies ran 49% longer, in tokens of reply, thinking not counted.

  • RAG and answering from documents: Claude Haiku 5.5 passed 5 of 5 and Gemini 3.8 Flash 5 of 5. Claude Haiku 5.5 answered 2.9× sooner at the median, 1.1 s against 3.1 s. Claude Haiku 5.5 cost 4.7× less, $0.0007 against $0.0033 for the 5 replies. Claude Haiku 5.5's replies ran 71% longer, in tokens of reply, thinking not counted.

  • Agents and tool use: Claude Haiku 5.5 passed 5 of 5 and Gemini 3.8 Flash 5 of 5. Claude Haiku 5.5 answered 4.2× sooner at the median, 1.0 s against 4.4 s. Claude Haiku 5.5 cost 6.3× less, $0.0007 against $0.0043 for the 5 replies. Claude Haiku 5.5's replies ran 34% longer, in tokens of reply, thinking not counted.

The 2 prompts only one of Claude and Gemini passed

Where one family's pick passed a prompt and the other's didn't, in each check's own words.

  • Announce a second bakery shop on LinkedIn (writing): Gemini 3.1 Pro (preview) passed and Claude Sonnet 5 didn't. Claude Sonnet 5: Graded 3.5 of 5 on average (lowest 3). Gemini 3.1 Pro (preview): Graded 4.0 of 5 on average (lowest 3).

  • Argue both sides of free buses (writing): Claude Sonnet 5 passed and Gemini 3.1 Pro (preview) didn't. Claude Sonnet 5: Graded 4.0 of 5 on average (lowest 4). Gemini 3.1 Pro (preview): Graded 3.7 of 5 on average (lowest 3).

Each Claude model against each Gemini model

Every Claude model against every Gemini model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.

  • Claude Fable 5.1 vs Gemini 3.1 Pro: 45 and 49 of 50; 4 prompts split them; Gemini 3.1 Pro's replies cost 1.9× less in all, and Claude Fable 5.1 answered sooner on 44, Gemini 3.1 Pro on 1.

  • Claude Fable 5.1 vs Gemini 3.8 Flash: 45 and 46 of 50; 7 prompts split them; Gemini 3.8 Flash's replies cost 15.6× less in all, and Gemini 3.8 Flash answered sooner on 37, Claude Fable 5.1 on 9.

  • Claude Opus 5.5 vs Gemini 3.1 Pro: 49 and 49 of 50; no prompt split them; their replies cost about the same in all, and Claude Opus 5.5 answered sooner on 45, Gemini 3.1 Pro on 2. Claude Opus 5.5 vs Gemini 3.1 Pro (preview).

  • Claude Opus 5.5 vs Gemini 3.8 Flash: 49 and 46 of 50; 5 prompts split them; Gemini 3.8 Flash's replies cost 8.3× less in all, and Gemini 3.8 Flash answered sooner on 26, Claude Opus 5.5 on 16.

  • Claude Sonnet 5.5 vs Gemini 3.1 Pro: 47 and 49 of 50; 2 prompts split them; Claude Sonnet 5.5's replies cost 2.8× less in all, and Claude Sonnet 5.5 answered sooner on 50, Gemini 3.1 Pro on 0. Claude Sonnet 5.5 vs Gemini 3.1 Pro (preview).

  • Claude Sonnet 5.5 vs Gemini 3.8 Flash: 47 and 46 of 50; 5 prompts split them; Gemini 3.8 Flash's replies cost 3.0× less in all, and Claude Sonnet 5.5 answered sooner on 47, Gemini 3.8 Flash on 1.

  • Claude Sonnet 5 vs Gemini 3.1 Pro: 46 and 49 of 50; 5 prompts split them; Claude Sonnet 5's replies cost 2.6× less in all, and Claude Sonnet 5 answered sooner on 50, Gemini 3.1 Pro on 0. Claude Sonnet 5 vs Gemini 3.1 Pro (preview).

  • Claude Sonnet 5 vs Gemini 3.8 Flash: 46 and 46 of 50; 6 prompts split them; Gemini 3.8 Flash's replies cost 3.2× less in all, and Claude Sonnet 5 answered sooner on 31, Gemini 3.8 Flash on 14. Claude Sonnet 5 vs Gemini 3.8 Flash.

  • Claude Haiku 4.5 vs Gemini 3.1 Pro: 43 and 49 of 50; 8 prompts split them; Claude Haiku 4.5's replies cost 7.2× less in all, and Claude Haiku 4.5 answered sooner on 50, Gemini 3.1 Pro on 0. Claude Haiku 4.5 vs Gemini 3.1 Pro (preview).

  • Claude Haiku 4.5 vs Gemini 3.8 Flash: 43 and 46 of 50; 5 prompts split them; Gemini 3.8 Flash's replies cost 1.2× less in all, and Claude Haiku 4.5 answered sooner on 44, Gemini 3.8 Flash on 3. Claude Haiku 4.5 vs Gemini 3.8 Flash.

  • Claude Haiku 5.5 vs Gemini 3.1 Pro: 45 and 49 of 50; 4 prompts split them; Claude Haiku 5.5's replies cost 45.4× less in all, and Claude Haiku 5.5 answered sooner on 50, Gemini 3.1 Pro on 0.

  • Claude Haiku 5.5 vs Gemini 3.8 Flash: 45 and 46 of 50; 5 prompts split them; Claude Haiku 5.5's replies cost 5.5× less in all, and Claude Haiku 5.5 answered sooner on 44, Gemini 3.8 Flash on 4.

One coding prompt, both replies

Claude Haiku 5.5 and Gemini 3.8 Flash on a coding prompt: both replies as they came.

A hard prompt: “Evaluate an arithmetic expression, no eval”, in full.

  • Claude Haiku 5.5

    Passed: All 15 tests passed.

    ```javascript
    function evaluate(expression) {
      const s = String(expression);
      const tokens = [];
      let i = 0;
      while (i < s.length) {
        const c = s[i];
        if (c === ' ' || c === '\t' || c === '\n' || c === '\r') { i++; continue; }
        if ((c >= '0' && c <= '9') || c === '.') {
          let j = i;
          while (j < s.length && ((s[j] >= '0' && s[j] <= '9') || s[j] === '.')) j++;
          const text = s.slice(i, j);
          if (!/^(\d+\.?\d*|\.\d+)$/.test(text)) throw new Error(`Invalid number "${text}"`);
          tokens.push({ type: 'num', value: Number(text) });
          i = j;
          continue;
        }
        if ('+-*/^()'.includes(c)) {
          tokens.push({ type: c });
          i++;
          continue;
        }
        throw new Error(`Unexpected character "${c}"`);
      }
    
      let pos = 0;
      const peek = () => tokens[pos] ? tokens[pos].type : null;
      const next = () => tokens[pos++];
    
      // expr := term (('+'|'-') term)*
      function parseExpr() {
        let value = parseTerm();
        while (peek() === '+' || peek() === '-') {
          const op = next().type;
          const rhs = parseTerm();
          value = op === '+' ? value + rhs : value - rhs;
        }
        return value;
      }
    
      // term := unary (('*'|'/') unary)*
      function parseTerm() {
        let value = parseUnary();
        while (peek() === '*' || peek() === '/') {
          const op = next().type;
          const rhs = parseUnary();
          value = op === '*' ? value * rhs : value / rhs;
        }
        return value;
      }
    
      // unary := '-' unary | power
      function parseUnary() {
        if (peek() === '-') {
          next();
          return -parseUnary();
        }
        return parsePower();
      }
    
      // power := primary ('^' unary)?   (right-associative; exponent may be signed)
      function parsePower() {
        const base = parsePrimary();
        if (peek() === '^') {
          next();
          const exp = parseUnary();
          return Math.pow(base, exp);
        }
        return base;
      }
    
      // primary := number | '(' expr ')'
      function parsePrimary() {
        const t = tokens[pos];
        if (!t) throw new Error('Unexpected end of expression');
        if (t.type === 'num') {
          pos++;
          return t.value;
        }
        if (t.type === '(') {
          pos++;
          const value = parseExpr();
          if (peek() !== ')') throw new Error('Expected ")"');
          pos++;
          return value;
        }
        throw new Error(`Unexpected token "${t.type}"`);
      }
    
      if (tokens.length === 0) throw new Error('Empty expression');
      const result = parseExpr();
      if (pos !== tokens.length) throw new Error(`Unexpected token "${tokens[pos].type}"`);
      return result;
    }
    ```

    763 tokens in, 1,074 out · 3.9 s · $0.0006 · 1 message on Pro · answered by anthropic/claude-haiku-5.5 via Anthropic ·

  • Gemini 3.8 Flash

    Passed: All 15 tests passed.

    ```javascript
    function evaluate(expression) {
      // Tokenize the expression
      const tokens = [];
      let i = 0;
      const n = expression.length;
    
      while (i < n) {
        const ch = expression[i];
    
        if (/\s/.test(ch)) {
          i++;
          continue;
        }
    
        if (/[0-9.]/.test(ch)) {
          let start = i;
          let dotCount = 0;
          while (i < n && /[0-9.]/.test(expression[i])) {
            if (expression[i] === '.') dotCount++;
            i++;
          }
          if (dotCount > 1) {
            throw new Error(`Invalid number format at index ${start}`);
          }
          const numStr = expression.slice(start, i);
          if (numStr === '.') {
            throw new Error(`Invalid single dot at index ${start}`);
          }
          tokens.push({ type: 'NUMBER', value: parseFloat(numStr) });
          continue;
        }
    
        if ('+-*/^()'.includes(ch)) {
          tokens.push({ type: ch, value: ch });
          i++;
          continue;
        }
    
        throw new Error(`Unexpected character '${ch}' at index ${i}`);
      }
    
      let pos = 0;
    
      function peek() {
        return tokens[pos];
      }
    
      function consume(expectedType) {
        const token = tokens[pos];
        if (!token || (expectedType && token.type !== expectedType)) {
          throw new Error(`Unexpected token at position ${pos}`);
        }
        pos++;
        return token;
      }
    
      // Grammar with precedence:
      // Expr    -> AddSub
      // AddSub  -> MulDiv (('+' | '-') MulDiv)*
      // MulDiv  -> Unary (('*' | '/') Unary)*
      // Unary   -> '-' Unary | Power
      // Power   -> Primary ('^' (Unary | Power))?   -- right associative, handles 2^-1 and -2^2
      // Primary -> NUMBER | '(' Expr ')'
    
      function parseExpr() {
        return parseAddSub();
      }
    
      function parseAddSub() {
        let left = parseMulDiv();
        while (pos < tokens.length && (tokens[pos].type === '+' || tokens[pos].type === '-')) {
          const op = consume().type;
          const right = parseMulDiv();
          left = op === '+' ? left + right : left - right;
        }
        return left;
      }
    
      function parseMulDiv() {
        let left = parseUnary();
        while (pos < tokens.length && (tokens[pos].type === '*' || tokens[pos].type === '/')) {
          const op = consume().type;
          const right = parseUnary();
          if (op === '/') {
            left = left / right;
          } else {
            left = left * right;
          }
        }
        return left;
      }
    
      function parseUnary() {
        if (pos < tokens.length && tokens[pos].type === '-') {
          consume('-');
          return -parseUnary();
        }
        if (pos < tokens.length && tokens[pos].type === '+') {
          consume('+');
          return parseUnary();
        }
        return parsePower();
      }
    
      function parsePower() {
        let left = parsePrimary();
        if (pos < tokens.length && tokens[pos].type === '^') {
          consume('^');
          // Right-hand side allows unary minus (e.g. 2^-1) or another power
          const right = parseUnary();
          return Math.pow(left, right);
        }
        return left;
      }
    
      function parsePrimary() {
        const token = peek();
        if (!token) {
          throw new Error('Unexpected end of input');
        }
    
        if (token.type === 'NUMBER') {
          consume('NUMBER');
          return token.value;
        }
    
        if (token.type === '(') {
          consume('(');
          const val = parseExpr();
          consume(')');
          return val;
        }
    
        throw new Error(`Unexpected token '${token.value}' at position ${pos}`);
      }
    
      if (tokens.length === 0) {
        throw new Error('Empty expression');
      }
    
      const result = parseExpr();
    
      if (pos < tokens.length) {
        throw new Error(`Unexpected token at position ${pos}`);
      }
    
      return result;
    }
    ```

    587 tokens in, 1,070 out · 6.7 s · $0.0045 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

Every model, every job

All 8 Claude and Gemini models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.

Every Claude and Gemini model across our test runs
ModelPassedHard onesCost per replyOn Pro
Claude Fable 5.1Anthropic45 of 5017 of 20$0.0200Up to 31 a month
Claude Opus 5.5Anthropic49 of 5019 of 20$0.0107Up to 62 a month
Claude Sonnet 5.5Anthropic47 of 5019 of 20$0.0038Up to 125 a month
Claude Sonnet 5Anthropic46 of 5019 of 20$0.0041Up to 125 a month
Claude Haiku 4.5Anthropic43 of 5015 of 20$0.0015Up to 250 a month
Claude Haiku 5.5Anthropic45 of 5019 of 20$0.00024Up to 60 a day
Gemini 3.1 Pro (preview)Google49 of 5019 of 20$0.0107Up to 125 a month
Gemini 3.8 FlashGoogle46 of 5018 of 20$0.0013Up to 250 a month

Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Their own subscriptions

Each company's own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.

llmwise Pro, $20 a month, has all 8 of these models in one chat, on one monthly allowance. On it: Claude Sonnet 5.5 up to 125 messages a month and Gemini 3.1 Pro (preview) up to 125.

The lineups at a glance

What follows from each model's facts in our catalog.

  • The lineups

    Claude: 6 models, Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 4.5, and Claude Haiku 5.5. Gemini: 2 models, Gemini 3.1 Pro (preview) and Gemini 3.8 Flash.

  • Price per message

    The least expensive Claude model is Claude Haiku 5.5 (60 messages a day on Pro); the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro).

  • Context window

    Claude goes up to 1M tokens (Claude Fable 5.1); Gemini up to 1.05M tokens (Gemini 3.1 Pro (preview)).

  • Images and PDFs

    Every model here reads images. Every model here takes a PDF as a whole file.

  • On the Free plan

    Free's one-time trial of 5 messages covers Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 4.5, Claude Haiku 5.5, Gemini 3.1 Pro (preview), and Gemini 3.8 Flash, and Claude Opus 5.5 for 1 message. Paid plans have every model, with messages every month.

Model by model

Two named models side by side, prompt by prompt, each with its messages on every plan.

More head-to-heads

Each of Claude and Gemini against the other families, every job from the same test runs.

Where your messages go

In llmwise, a message to Claude or Gemini goes to the model's maker, or through OpenRouter when llmwise can't reach the maker directly. The Privacy Policy has the details.

Bar chart: Prompts passed in our test runs, all 10 jobs. Gemini 3.1 Pro (preview): 49 of 50; Claude Opus 5.5: 49 of 50; Claude Sonnet 5.5: 47 of 50; Claude Sonnet 5: 46 of 50; Gemini 3.8 Flash: 46 of 50; Claude Haiku 5.5: 45 of 50; Claude Fable 5.1: 45 of 50; Claude Haiku 4.5: 43 of 50.
Our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026: the same prompts for every model, each reply checked the same way.

Questions

Which is better, Claude or Gemini?

In our test runs on October 7, 2026, the same 50 prompts across 10 jobs: Claude's 6 models passed 275 of 300 replies and Gemini's 2 models passed 95 of 100 replies. Job by job, both families' best models shared the top result on 10 of the 10 jobs.

Which is cheaper, Claude or Gemini?

In llmwise, the least expensive Claude model is Claude Haiku 5.5 (60 messages a day on Pro), and the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro). At API list prices (October 2026), a typical message of 4,000 tokens in and 700 out costs $0.0008 on Claude Haiku 5.5 and $0.0056 on Gemini 3.8 Flash.

Can I use Claude and Gemini in the same chat?

Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.