Skip to content

Gemini · Coding

Gemini for coding

llmwise has 2 of Google's Gemini models, from Gemini 3.8 Flash (up to 250 messages a month on Pro) to Gemini 3.1 Pro (preview) (up to 125 messages a month). We ran the same coding prompts on every one and published every reply: which Gemini model to use, from the results, what each costs per message, and how to get more out of it.

Based on 10 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

In our test runs on September 27, 2026, all 2 models passed 5 of 5 coding prompts, so these prompts don't pick one for hard problems. For value, Gemini 3.8 Flash (5 of 5), 250 a month on Pro. It's also our pick for everyday.

Our picks for coding

  • Hard problems

    Shared by 2 models

    All 2 models passed 5 of 5, both hard ones: Gemini 3.1 Pro and Gemini 3.8 Flash. These prompts don't tell them apart, so they share the pick.

  • Best value

    Gemini 3.8 Flash

    Passed 5 of 5 coding prompts, with 250 a month on Pro.

  • Everyday

    Gemini 3.8 Flash

    Passed 5 of 5 coding prompts, at 250 a month on Pro. The same model as best value.

These picks aren't our opinion: they're what the results below give, by these rules, among the Gemini models in llmwise. They change when the results do.

  • Hard problems: the model that passed the most prompts and, of those, the most hard ones. Models level on both share the pick: the prompts don't tell them apart, so we don't break the tie by price or by name.
  • Best value: among the models that draw on the monthly allowance, the one with the most messages on Pro that passed no more than one prompt fewer than the top model. Ties go to the one that passed more, then to the lower cost per reply.
  • Everyday: among the cheapest models on the page (the everyday models, which come from the daily count, when the page has any), the one that passed the most. Ties go to the one that passed more of the hard prompts, then to the lower cost per reply.

Our coding test runs, model by model

How each model did on our 5 coding prompts, what each reply counted as on Pro, and what it cost to run.

Our coding test runs
ModelPassedHard onesMessages used on ProCost per replyTime per reply
Gemini 3.1 Pro (preview)Google5 of 52 of 21 each, of 125 a month on Pro$0.019513.2 s
Gemini 3.8 FlashGoogle5 of 52 of 21 each, of 250 a month on Pro$0.00204.4 s
Passed: out of the job's 5 prompts, how many replies passed their check (a reply the provider failed to give isn't counted). Messages used: what each reply counted as on Pro, and Pro's count on the model. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Time: from sending to the whole reply.

One prompt, side by side

Each model's reply to “Evaluate an arithmetic expression, no eval”, as it came, with whether it passed and what it cost.

The prompt: Evaluate an arithmetic expression, no eval (hard)
Write a JavaScript function evaluate(expression) that computes an arithmetic expression given as a string and returns the number.
It supports numbers like 3, 0.5 and 12.25; the operators + - * / and ^ (power); parentheses; unary minus; and spaces anywhere.
- ^ binds tighter than unary minus and is right-associative, so -2^2 is -4 and 2^3^2 is 512. A unary minus may follow ^, as in 2^-1, which is 0.5.
- * and / bind tighter than + and -. Otherwise operators of the same level go left to right, so 8/4/2 is 1.
- Throw an Error for input that isn't a valid expression, such as "2 +", "(1" or "1 2".
Don't use eval, Function or any library.
Reply with the whole function in one ```javascript code block: plain JavaScript for Node.js 22, no imports, no TypeScript.
  • Gemini 3.1 Pro

    Passed: All 15 tests passed.

    ```javascript
    function evaluate(expression) {
        let pos = 0;
    
        function peek() {
            while (pos < expression.length && expression[pos] === ' ') {
                pos++;
            }
            return pos < expression.length ? expression[pos] : null;
        }
    
        function parseExpr() {
            let left = parseTerm();
            while (true) {
                let char = peek();
                if (char === '+' || char === '-') {
                    pos++;
                    let right = parseTerm();
                    if (char === '+') left += right;
                    else left -= right;
                } else {
                    break;
                }
            }
            return left;
        }
    
        function parseTerm() {
            let left = parseFactor();
            while (true) {
                let char = peek();
                if (char === '*' || char === '/') {
                    pos++;
                    let right = parseFactor();
                    if (char === '*') left *= right;
                    else left /= right;
                } else {
                    break;
                }
            }
            return left;
        }
    
        function parseFactor() {
            let char = peek();
            if (char === '-') {
                pos++;
                return -parseFactor();
            }
            return parsePower();
        }
    
        function parsePower() {
            let left = parsePrimary();
            let char = peek();
            if (char === '^') {
                pos++;
                let right = parseFactor();
                return Math.pow(left, right);
            }
            return left;
        }
    
        function parsePrimary() {
            let char = peek();
            if (char === null) {
                throw new Error("Unexpected end of expression");
            }
    
            if (char === '(') {
                pos++;
                let val = parseExpr();
                if (peek() !== ')') {
                    throw new Error("Missing closing parenthesis");
                }
                pos++;
                return val;
            }
    
            let match = expression.slice(pos).match(/^\d+(\.\d+)?/);
            if (!match) {
                throw new Error(`Invalid character at position ${pos}`);
            }
            pos += match[0].length;
            return parseFloat(match[0]);
        }
    
        let result = parseExpr();
        if (peek() !== null) {
            throw new Error(`Unexpected character at position ${pos}`);
        }
        
        return result;
    }
    ```

    587 tokens in, 2,831 out (2,262 of them reasoning) · 20.6 s · $0.0351 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  • Gemini 3.8 Flash

    Passed: All 15 tests passed.

    ```javascript
    function evaluate(expression) {
      // Tokenize the expression
      const tokens = [];
      let i = 0;
      const n = expression.length;
    
      while (i < n) {
        const ch = expression[i];
    
        if (/\s/.test(ch)) {
          i++;
          continue;
        }
    
        if (/[0-9.]/.test(ch)) {
          let start = i;
          let dotCount = 0;
          while (i < n && /[0-9.]/.test(expression[i])) {
            if (expression[i] === '.') dotCount++;
            i++;
          }
          if (dotCount > 1) {
            throw new Error(`Invalid number format at index ${start}`);
          }
          const numStr = expression.slice(start, i);
          if (numStr === '.') {
            throw new Error(`Invalid single dot at index ${start}`);
          }
          tokens.push({ type: 'NUMBER', value: parseFloat(numStr) });
          continue;
        }
    
        if ('+-*/^()'.includes(ch)) {
          tokens.push({ type: ch, value: ch });
          i++;
          continue;
        }
    
        throw new Error(`Unexpected character '${ch}' at index ${i}`);
      }
    
      let pos = 0;
    
      function peek() {
        return tokens[pos];
      }
    
      function consume(expectedType) {
        const token = tokens[pos];
        if (!token || (expectedType && token.type !== expectedType)) {
          throw new Error(`Unexpected token at position ${pos}`);
        }
        pos++;
        return token;
      }
    
      // Grammar with precedence:
      // Expr    -> AddSub
      // AddSub  -> MulDiv (('+' | '-') MulDiv)*
      // MulDiv  -> Unary (('*' | '/') Unary)*
      // Unary   -> '-' Unary | Power
      // Power   -> Primary ('^' (Unary | Power))?   -- right associative, handles 2^-1 and -2^2
      // Primary -> NUMBER | '(' Expr ')'
    
      function parseExpr() {
        return parseAddSub();
      }
    
      function parseAddSub() {
        let left = parseMulDiv();
        while (pos < tokens.length && (tokens[pos].type === '+' || tokens[pos].type === '-')) {
          const op = consume().type;
          const right = parseMulDiv();
          left = op === '+' ? left + right : left - right;
        }
        return left;
      }
    
      function parseMulDiv() {
        let left = parseUnary();
        while (pos < tokens.length && (tokens[pos].type === '*' || tokens[pos].type === '/')) {
          const op = consume().type;
          const right = parseUnary();
          if (op === '/') {
            left = left / right;
          } else {
            left = left * right;
          }
        }
        return left;
      }
    
      function parseUnary() {
        if (pos < tokens.length && tokens[pos].type === '-') {
          consume('-');
          return -parseUnary();
        }
        if (pos < tokens.length && tokens[pos].type === '+') {
          consume('+');
          return parseUnary();
        }
        return parsePower();
      }
    
      function parsePower() {
        let left = parsePrimary();
        if (pos < tokens.length && tokens[pos].type === '^') {
          consume('^');
          // Right-hand side allows unary minus (e.g. 2^-1) or another power
          const right = parseUnary();
          return Math.pow(left, right);
        }
        return left;
      }
    
      function parsePrimary() {
        const token = peek();
        if (!token) {
          throw new Error('Unexpected end of input');
        }
    
        if (token.type === 'NUMBER') {
          consume('NUMBER');
          return token.value;
        }
    
        if (token.type === '(') {
          consume('(');
          const val = parseExpr();
          consume(')');
          return val;
        }
    
        throw new Error(`Unexpected token '${token.value}' at position ${pos}`);
      }
    
      if (tokens.length === 0) {
        throw new Error('Empty expression');
      }
    
      const result = parseExpr();
    
      if (pos < tokens.length) {
        throw new Error(`Unexpected token at position ${pos}`);
      }
    
      return result;
    }
    ```

    587 tokens in, 1,070 out · 6.7 s · $0.0045 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

The coding prompts, and how they're scored

Tests. Automatic. The function runs against the prompt's tests in a separate Node.js process with a time limit and no file, network or child-process access; it passes when every test passes.

Each prompt was sent the way llmwise sends a message in a side-by-side comparison, which offers no tools: the app's own system prompt, the model's own settings, and Pro's reply size limit (8,000 tokens). Read all five coding prompts and how every reply was scored.

Prompts like these to try yourself

Gemini in llmwise

Gemini models in llmwise
ModelOn ProOn FreeContext windowImagesPDFsReasoningAPI price per 1M, in / out
Gemini 3.1 Pro (preview)Google125/mo on ProYes1.05M tokensYesWhole fileYes$2.00 / $12.00
Gemini 3.8 FlashGoogle250/mo on ProYes1.05M tokensYesWhole fileYes$0.75 / $3.75
Each badge is how many messages Pro gets on the model: a month’s, or a day’s on an everyday model. Free is a one-time trial of 5 messages on the models marked. “Text only” models get the text of a PDF, not the file. API prices are the per-token prices in our model catalog as of October 2026 (Google: Google's list price). In llmwise you pay per message, not per token. Gemini 3.1 Pro (preview): the standard rate, for prompts up to 200K tokens. Gemini 3.8 Flash: an introductory price, through December 31, 2026.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

What matters for coding

  • Room for your code

    The more of your code a model can take in at once, the fewer details it has to guess. Context windows are listed below.

  • Reasoning on hard problems

    Models that reason before answering tend to do better on problems with several moving parts, like a bug that crosses files.

  • Running the code

    Code that has been run beats code that looks right. A model that can run and test its own code catches its own mistakes.

  • Cost per try

    Coding means many small tries. A cheap model for quick questions and a stronger one for the hard ones keeps the cost down.

Gemini for coding in llmwise

  • Code opens as a document

    Ask for a script or a component and it opens beside the chat as a code document that keeps every version, so you can step back to an earlier one.

  • Run code (paid plans)

    On a paid plan the model can run Python or Node.js in an isolated sandbox once you approve it, read the output and fix what failed. A run counts as 1 Claude Haiku 4.5 message and stops after 60 seconds.

  • HTML apps run in the chat

    Ask for a calculator, a small tool or a page and it runs live as an HTML app in the side panel.

  • Switch models mid-chat

    Start on a cheaper model; if the answer isn't good enough, switch models in the same chat. The next model sees the whole conversation, so you don't paste anything twice.

In llmwise, a message to Gemini goes to its maker, Google, or through OpenRouter when llmwise can't reach the maker directly. See the Privacy Policy.

Tips

  • Paste the exact error and the code that produced it, not just a description.

  • Say which language, version and libraries you use.

  • For a big change, ask for a plan first, then one step at a time.

  • Pick the Coder persona for code-first answers that explain their trade-offs.

  • If a model gets stuck, switch to another in the same chat: it sees the code so far.

Questions

Is Gemini good for coding?

In our test runs on September 27, 2026, Gemini 3.1 Pro (preview) passed 5 of 5, Gemini 3.8 Flash passed 5 of 5 of our coding prompts, against a best result of 5 of 5 among all 19 models. Every reply is published on this page and the methods page, so you can judge them yourself.

Which Gemini model should I use for coding?

Start with Gemini 3.8 Flash for most day-to-day coding: writing functions, explaining code, reviewing a change, and move up to Gemini 3.1 Pro (preview) for most day-to-day coding: writing functions, explaining code, reviewing a change.

Can I use Gemini for coding for free?

Yes: Gemini 3.1 Pro (preview) and Gemini 3.8 Flash are on the Free plan (a one-time trial of 5 messages).

Can llmwise run the code it writes?

On a paid plan, yes. The model writes Python or Node.js, you approve the run, and it executes in an isolated sandbox with no internet access except package registries, stopping after 60 seconds. It reads the output and can fix what failed. Each run counts as 1 Claude Haiku 4.5 message.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.