Skip to content

Best AI · Coding

The best AI for coding

In our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026, 17 of the 19 models passed 5 of 5 coding prompts, so these prompts don't name one best model. For best value, GLM 5.3 (5 of 5). Our picks below follow fixed rules, beside each model's price per message, and every prompt and reply is published.

Based on 95 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

In our test runs on October 9, 2026, 17 models passed 5 of 5 coding prompts, so these prompts don't pick one for hard problems. For value, GLM 5.3 (5 of 5), 250 a month on Pro; for everyday coding, GLM 5.3 Flash (5 of 5), from the daily count.

Our picks for coding

  • Hard problems

    Shared by 17 models

    17 models passed 5 of 5, both hard ones: Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5 and 14 more. These prompts don't tell them apart, so they share the pick.

  • Best value

    GLM 5.3

    Passed 5 of 5 coding prompts, with 250 a month on Pro.

  • Everyday

    GLM 5.3 Flash

    Passed 5 of 5 coding prompts; an everyday model, so its messages come from the daily count (60 a day on Pro), not the monthly allowance.

These picks aren't our opinion: they're what the results below give, by these rules, among all 19 models in llmwise. They change when the results do.

  • Hard problems: the model that passed the most prompts and, of those, the most hard ones. Models level on both share the pick: the prompts don't tell them apart, so we don't break the tie by price or by name.
  • Best value: among the models that draw on the monthly allowance, the one with the most messages on Pro that passed no more than one prompt fewer than the top model. Ties go to the one that passed more, then to the lower cost per reply.
  • Everyday: among the cheapest models on the page (the everyday models, which come from the daily count, when the page has any), the one that passed the most. Ties go to the one that passed more of the hard prompts, then to the lower cost per reply.

Our coding test runs, model by model

How each model did on our 5 coding prompts, what each reply counted as on Pro, and what it cost to run.

Our coding test runs
ModelPassedHard onesMessages used on ProCost per replyTime per reply
Claude Fable 5.1Anthropic5 of 52 of 21 each, of 31 a month on Pro$0.04259.1 s
Claude Opus 5.5Anthropic5 of 52 of 21 each, of 62 a month on Pro$0.01718.3 s
Claude Sonnet 5.5Anthropic5 of 52 of 21 each, of 125 a month on Pro$0.00633.0 s
Claude Sonnet 5Anthropic5 of 52 of 21 each, of 125 a month on Pro$0.01028.0 s
Claude Haiku 5.5Anthropic5 of 52 of 21 each, of 60 a day on Pro$0.00054.0 s
Claude Haiku 4.5Anthropic3 of 50 of 21 each, of 250 a month on Pro$0.00273.4 s
GPT-6 AstraOpenAI5 of 52 of 21 each, of 31 a month on Pro$0.02236.8 s
GPT-6.1 SolOpenAI5 of 52 of 21 each, of 125 a month on Pro$0.00234.9 s
GPT-6 SolOpenAI5 of 52 of 21 each, of 125 a month on Pro$0.00525.7 s
GPT-6 LunaOpenAI5 of 52 of 21 each, of 60 a day on Pro$0.00034.9 s
Gemini 3.1 Pro (preview)Google5 of 52 of 21 each, of 125 a month on Pro$0.019513.2 s
Gemini 3.8 FlashGoogle5 of 52 of 21 each, of 250 a month on Pro$0.00204.4 s
DeepSeek V4.1 FlashDeepSeek5 of 52 of 21 each, of 60 a day on Pro$0.00183.9 s
DeepSeek V4 ProDeepSeek5 of 52 of 21 each, of 250 a month on Pro$0.024079.6 s
Grok 4.7xAI5 of 52 of 21 each, of 250 a month on Pro$0.024646.9 s
Kimi K3Moonshot5 of 52 of 21 each, of 125 a month on Pro$0.00635.2 s
GLM 5.3Z.ai5 of 52 of 21 each, of 250 a month on Pro$0.00153.5 s
GLM 5.3 FlashZ.ai5 of 52 of 21 each, of 60 a day on Pro$0.00038.0 s
Mistral Large 4Mistral4 of 51 of 21 each, of 250 a month on Pro$0.007341.0 s
Passed: out of the job's 5 prompts, how many replies passed their check (a reply the provider failed to give isn't counted). Messages used: what each reply counted as on Pro, and Pro's count on the model. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Time: from sending to the whole reply.

One prompt, side by side

Each model's reply to “Evaluate an arithmetic expression, no eval”, as it came, with whether it passed and what it cost.

The prompt: Evaluate an arithmetic expression, no eval (hard)
Write a JavaScript function evaluate(expression) that computes an arithmetic expression given as a string and returns the number.
It supports numbers like 3, 0.5 and 12.25; the operators + - * / and ^ (power); parentheses; unary minus; and spaces anywhere.
- ^ binds tighter than unary minus and is right-associative, so -2^2 is -4 and 2^3^2 is 512. A unary minus may follow ^, as in 2^-1, which is 0.5.
- * and / bind tighter than + and -. Otherwise operators of the same level go left to right, so 8/4/2 is 1.
- Throw an Error for input that isn't a valid expression, such as "2 +", "(1" or "1 2".
Don't use eval, Function or any library.
Reply with the whole function in one ```javascript code block: plain JavaScript for Node.js 22, no imports, no TypeScript.
  • Claude Fable 5.1

    Passed: All 15 tests passed.

    ```javascript
    function evaluate(expression) {
      if (typeof expression !== "string") throw new Error("Expression must be a string");
    
      // ---- Tokenizer ----
      const tokens = [];
      let i = 0;
      while (i < expression.length) {
        const ch = expression[i];
        if (ch === " " || ch === "\t" || ch === "\n" || ch === "\r") { i++; continue; }
        if (ch >= "0" && ch <= "9") {
          let start = i;
          while (i < expression.length && expression[i] >= "0" && expression[i] <= "9") i++;
          if (expression[i] === ".") {
            i++;
            if (!(expression[i] >= "0" && expression[i] <= "9")) {
              throw new Error(`Invalid number at position ${start}`);
            }
            while (i < expression.length && expression[i] >= "0" && expression[i] <= "9") i++;
          }
          tokens.push({ type: "num", value: parseFloat(expression.slice(start, i)), pos: start });
          continue;
        }
        if ("+-*/^()".includes(ch)) {
          tokens.push({ type: ch, pos: i });
          i++;
          continue;
        }
        throw new Error(`Unexpected character '${ch}' at position ${i}`);
      }
      tokens.push({ type: "end", pos: expression.length });
    
      // ---- Parser ----
      let p = 0;
      const peek = () => tokens[p];
      const next = () => tokens[p++];
      const fail = (tok) =>
        new Error(
          tok.type === "end"
            ? "Unexpected end of expression"
            : `Unexpected token '${tok.type === "num" ? tok.value : tok.type}' at position ${tok.pos}`
        );
    
      // expr := term (('+'|'-') term)*
      function parseExpr() {
        let value = parseTerm();
        while (peek().type === "+" || peek().type === "-") {
          const op = next().type;
          const rhs = parseTerm();
          value = op === "+" ? value + rhs : value - rhs;
        }
        return value;
      }
    
      // term := unary (('*'|'/') unary)*
      function parseTerm() {
        let value = parseUnary();
        while (peek().type === "*" || peek().type === "/") {
          const op = next().type;
          const rhs = parseUnary();
          value = op === "*" ? value * rhs : value / rhs;
        }
        return value;
      }
    
      // unary := '-' unary | power
      function parseUnary() {
        if (peek().type === "-") {
          next();
          return -parseUnary();
        }
        return parsePower();
      }
    
      // power := atom ('^' unary)?   (right-associative; exponent may carry unary minus)
      function parsePower() {
        const base = parseAtom();
        if (peek().type === "^") {
          next();
          const exponent = parseUnary();
          return Math.pow(base, exponent);
        }
        return base;
      }
    
      // atom := number | '(' expr ')'
      function parseAtom() {
        const tok = next();
        if (tok.type === "num") return tok.value;
        if (tok.type === "(") {
          const value = parseExpr();
          const close = next();
          if (close.type !== ")") throw fail(close);
          return value;
        }
        throw fail(tok);
      }
    
      const result = parseExpr();
      if (peek().type !== "end") throw fail(peek());
      return result;
    }
    ```

    762 tokens in, 1,561 out (387 of them reasoning) · 15.7 s · $0.0874 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

  • GLM 5.3

    Passed: All 15 tests passed.

    ```javascript
    function evaluate(expression) {
      if (typeof expression !== 'string') throw new Error('Invalid expression');
      const src = expression;
      let pos = 0;
    
      function skipSpaces() {
        while (pos < src.length && /\s/.test(src[pos])) pos++;
      }
    
      function parseExpr() {
        let left = parseTerm();
        for (;;) {
          skipSpaces();
          if (src[pos] === '+') { pos++; left += parseTerm(); }
          else if (src[pos] === '-') { pos++; left -= parseTerm(); }
          else return left;
        }
      }
    
      function parseTerm() {
        let left = parseUnary();
        for (;;) {
          skipSpaces();
          if (src[pos] === '*') { pos++; left *= parseUnary(); }
          else if (src[pos] === '/') { pos++; left /= parseUnary(); }
          else return left;
        }
      }
    
      function parseUnary() {
        skipSpaces();
        if (src[pos] === '-') { pos++; return -parseUnary(); }
        return parsePower();
      }
    
      function parsePower() {
        const base = parsePrimary();
        skipSpaces();
        if (src[pos] === '^') {
          pos++;
          return Math.pow(base, parseUnary()); // right-assoc, exponent may be unary
        }
        return base;
      }
    
      function parsePrimary() {
        skipSpaces();
        if (src[pos] === '(') {
          pos++;
          const value = parseExpr();
          skipSpaces();
          if (src[pos] !== ')') throw new Error('Invalid expression: expected )');
          pos++;
          return value;
        }
        if (src[pos] === '-') return parseUnary(); // handled here too for safety
        const start = pos;
        while (pos < src.length && /[0-9.]/.test(src[pos])) pos++;
        if (pos === start) throw new Error('Invalid expression: expected number');
        const text = src.slice(start, pos);
        if (!/^\d+(\.\d+)?$|^\.\d+$/.test(text)) throw new Error('Invalid number: ' + text);
        return parseFloat(text);
      }
    
      const result = parseExpr();
      skipSpaces();
      if (pos !== src.length) throw new Error('Invalid expression: unexpected trailing input');
      return result;
    }
    ```

    562 tokens in, 594 out (85 of them reasoning) · 12.5 s · $0.0017 · 1 message on Pro · answered by z-ai/glm-5.3 via Reka ·

  • GLM 5.3 Flash

    Passed: All 15 tests passed.

    ```javascript
    function evaluate(expression) {
      let i = 0;
      const s = expression;
    
      function skipSpaces() {
        while (i < s.length && s[i] === ' ') i++;
      }
    
      function parseExpr() {
        let value = parseTerm();
        for (;;) {
          skipSpaces();
          if (s[i] === '+') { i++; value += parseTerm(); }
          else if (s[i] === '-') { i++; value -= parseTerm(); }
          else return value;
        }
      }
    
      function parseTerm() {
        let value = parseUnary();
        for (;;) {
          skipSpaces();
          if (s[i] === '*') { i++; value *= parseUnary(); }
          else if (s[i] === '/') { i++; value /= parseUnary(); }
          else return value;
        }
      }
    
      function parseUnary() {
        skipSpaces();
        if (s[i] === '-') { i++; return -parseUnary(); }
        return parsePower();
      }
    
      function parsePower() {
        const base = parsePrimary();
        skipSpaces();
        if (s[i] === '^') {
          i++;
          return Math.pow(base, parseUnary()); // right-assoc, allows 2^-1
        }
        return base;
      }
    
      function parsePrimary() {
        skipSpaces();
        if (s[i] === '(') {
          i++;
          const value = parseExpr();
          skipSpaces();
          if (s[i] !== ')') throw new Error("Expected ')'");
          i++;
          return value;
        }
        const start = i;
        while (i < s.length && s[i] >= '0' && s[i] <= '9') i++;
        if (i < s.length && s[i] === '.') {
          i++;
          while (i < s.length && s[i] >= '0' && s[i] <= '9') i++;
        }
        if (i === start || (i === start + 1 && s[start] === '.'))
          throw new Error('Invalid expression');
        return parseFloat(s.slice(start, i));
      }
    
      const result = parseExpr();
      skipSpaces();
      if (i !== s.length) throw new Error('Invalid expression');
      return result;
    }
    ```

    563 tokens in, 595 out (118 of them reasoning) · 11.1 s · $0.0004 · 1 message on Pro · answered by z-ai/glm-5.3-flash via AtlasCloud ·

The coding prompts, and how they're scored

Tests. Automatic. The function runs against the prompt's tests in a separate Node.js process with a time limit and no file, network or child-process access; it passes when every test passes.

Each prompt was sent the way llmwise sends a message in a side-by-side comparison, which offers no tools: the app's own system prompt, the model's own settings, and Pro's reply size limit (8,000 tokens). Read all five coding prompts and how every reply was scored.

Prompts like these to try yourself

What can five coding prompts tell you?

Less than a leaderboard suggests. 17 of the 19 models passed all five of our coding prompts, from GLM 5.3 Flash at $0.0003 a reply to Claude Fable 5.1 at $0.0425: about 168 times the price for the same result.

So for everyday code (a function, a query, a fix you can test), the price per message matters more than the ranking. Our prompts can't separate the top models on a large codebase or a subtle bug. For that, send your own code to two of them in one chat and compare.

Context windows decide how much code fits in one message, and each maker states its own: Anthropic's models overview and OpenAI's GPT-6 Astra page.

What matters for coding

  • Room for your code

    The more of your code a model can take in at once, the fewer details it has to guess. Context windows are listed below.

  • Reasoning on hard problems

    Models that reason before answering tend to do better on problems with several moving parts, like a bug that crosses files.

  • Running the code

    Code that has been run beats code that looks right. A model that can run and test its own code catches its own mistakes.

  • Cost per try

    Coding means many small tries. A cheap model for quick questions and a stronger one for the hard ones keeps the cost down.

Every model at a glance

Every model in llmwise
ModelOn ProOn FreeContext windowImagesPDFsReasoningAPI price per 1M, in / out
Claude Fable 5.1Anthropic31/mo on ProNo1M tokensYesWhole fileYes$10.00 / $50.00
Claude Opus 5.5Anthropic62/mo on Pro1 message1M tokensYesWhole fileYes$4.00 / $20.00
Claude Sonnet 5.5Anthropic125/mo on ProYes1M tokensYesWhole fileYes$2.00 / $10.00
Claude Sonnet 5Anthropic125/mo on ProYes1M tokensYesWhole fileYes$2.00 / $10.00
Claude Haiku 5.5Anthropic60/day on ProYes1M tokensYesWhole fileYes$0.10 / $0.50
Claude Haiku 4.5Anthropic250/mo on ProYes200K tokensYesWhole fileNo$1.00 / $5.00
GPT-6 AstraOpenAI31/mo on ProNo1.05M tokensYesWhole fileYes$10.00 / $50.00
GPT-6.1 SolOpenAI125/mo on ProYes1.05M tokensYesWhole fileYes$2.00 / $10.00
GPT-6 SolOpenAI125/mo on ProYes1.05M tokensYesWhole fileYes$2.00 / $10.00
GPT-6 LunaOpenAI60/day on ProYes1.05M tokensYesWhole fileYes$0.10 / $0.50
Gemini 3.1 Pro (preview)Google125/mo on ProYes1.05M tokensYesWhole fileYes$2.00 / $12.00
Gemini 3.8 FlashGoogle250/mo on ProYes1.05M tokensYesWhole fileYes$0.75 / $3.75
DeepSeek V4.1 FlashDeepSeek60/day on ProYes1.05M tokensYesText onlyYes$0.30 / $1.20
DeepSeek V4 ProDeepSeek250/mo on ProYes1.05M tokensNoText onlyYes$0.40 / $4.00
Grok 4.7xAI250/mo on ProYes500K tokensYesWhole fileYes$2.00 / $6.00
Kimi K3Moonshot125/mo on ProYes1.05M tokensYesText onlyYes$3.00 / $15.00
GLM 5.3Z.ai250/mo on ProYes1.05M tokensNoText onlyYes$1.40 / $4.40
GLM 5.3 FlashZ.ai60/day on ProYes1.05M tokensYesText onlyYes$0.15 / $0.50
Mistral Large 4Mistral250/mo on ProYes1.05M tokensYesText onlyYes$0.68 / $2.09
Each badge is how many messages Pro gets on the model: a month’s, or a day’s on an everyday model. Free is a one-time trial of 5 messages on the models marked. “Text only” models get the text of a PDF, not the file. API prices are the per-token prices in our model catalog as of October 2026 (Anthropic: Anthropic's list price; OpenAI: OpenAI's list price; Google: Google's list price; DeepSeek: the price of the OpenRouter endpoints llmwise uses, not DeepSeek's own API; xAI: xAI's price, served through OpenRouter; Moonshot: Moonshot's list price; Z.ai: Z.ai's list price; Mistral: Mistral's price, served through OpenRouter). In llmwise you pay per message, not per token. Claude Haiku 5.5: the rate for prompts up to 100K tokens; $0.50 / $2.50 a million over that. Gemini 3.1 Pro (preview): the standard rate, for prompts up to 200K tokens. Gemini 3.8 Flash: an introductory price, through December 31, 2026. Grok 4.7: xAI charges more for very long prompts. Mistral Large 4: a sale price, half its list price of $1.36 / $4.18.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Coding in llmwise

  • Code opens as a document

    Ask for a script or a component and it opens beside the chat as a code document that keeps every version, so you can step back to an earlier one.

  • Run code (paid plans)

    On a paid plan the model can run Python or Node.js in an isolated sandbox once you approve it, read the output and fix what failed. A run counts as 1 Claude Haiku 4.5 message and stops after 60 seconds.

  • HTML apps run in the chat

    Ask for a calculator, a small tool or a page and it runs live as an HTML app in the side panel.

  • Switch models mid-chat

    Start on a cheaper model; if the answer isn't good enough, switch models in the same chat. The next model sees the whole conversation, so you don't paste anything twice.

Tips

  • Paste the exact error and the code that produced it, not just a description.

  • Say which language, version and libraries you use.

  • For a big change, ask for a plan first, then one step at a time.

  • Pick the Coder persona for code-first answers that explain their trade-offs.

  • If a model gets stuck, switch to another in the same chat: it sees the code so far.

Bar chart: Prompts passed in our test runs, coding. GLM 5.3 Flash: 5 of 5; GPT-6 Luna: 5 of 5; Claude Haiku 5.5: 5 of 5; DeepSeek V4.1 Flash: 5 of 5; GLM 5.3: 5 of 5; Gemini 3.8 Flash: 5 of 5; DeepSeek V4 Pro: 5 of 5; Grok 4.7: 5 of 5; GPT-6.1 Sol: 5 of 5; GPT-6 Sol: 5 of 5; Claude Sonnet 5.5: 5 of 5; Kimi K3: 5 of 5; Claude Sonnet 5: 5 of 5; Gemini 3.1 Pro (preview): 5 of 5; Claude Opus 5.5: 5 of 5; GPT-6 Astra: 5 of 5; Claude Fable 5.1: 5 of 5; Mistral Large 4: 4 of 5; Claude Haiku 4.5: 3 of 5.
Our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026: the same prompts for every model, each reply checked the same way.

Coding with an LLM: questions

What is the best AI for coding?

In our test runs on October 9, 2026, 17 of the 19 models passed 5 of 5 coding prompts, so these prompts don't name one best model. For best value, GLM 5.3 (5 of 5). For everyday, GLM 5.3 Flash (5 of 5). Every prompt and reply is published, so you can check them, and the picks follow fixed rules.

How did you test the models for coding?

We sent the same coding prompts to every model through llmwise's own pipeline and checked each reply the same way. The prompts, the replies, how each was scored and the grader are all published on the methods page.

Can I try these models for coding for free?

Yes, to try: Free is a one-time trial of 5 messages on every model but Claude Fable 5.1 and GPT-6 Astra (one of them can be on Claude Opus 5.5).

Can llmwise run the code it writes?

On a paid plan, yes. The model writes Python or Node.js, you approve the run, and it executes in an isolated sandbox with no internet access except package registries, stopping after 60 seconds. It reads the output and can fix what failed. Each run counts as 1 Claude Haiku 4.5 message.

Does it connect to my repository or editor?

No. llmwise is a chat app: paste the code you're asking about. It doesn't read your repository or plug into an editor. On a paid plan you can connect MCP servers that give the model extra tools.

Try your own code, not our five prompts

Paste the bug you're on into Claude Sonnet 5 (up to 125 a month on Pro), then ask Claude Fable 5.1 (up to 31 a month) what it would change. Both see the same chat.