Skip to content

Gemini · Math

Gemini for math

llmwise has 2 of Google's Gemini models, from Gemini 3.8 Flash (up to 250 messages a month on Pro) to Gemini 3.1 Pro (preview) (up to 125 messages a month). We ran the same math prompts on every one and published every reply: which Gemini model to use, from the results, what each costs per message, and how to get more out of it.

Based on 10 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

In our test runs on September 27, 2026, all 2 models passed 5 of 5 math prompts, so these prompts don't pick one for hard problems. For value, Gemini 3.8 Flash (5 of 5), 250 a month on Pro. It's also our pick for everyday.

Our picks for math

  • Hard problems

    Shared by 2 models

    All 2 models passed 5 of 5, both hard ones: Gemini 3.1 Pro and Gemini 3.8 Flash. These prompts don't tell them apart, so they share the pick.

  • Best value

    Gemini 3.8 Flash

    Passed 5 of 5 math prompts, with 250 a month on Pro.

  • Everyday

    Gemini 3.8 Flash

    Passed 5 of 5 math prompts, at 250 a month on Pro. The same model as best value.

These picks aren't our opinion: they're what the results below give, by these rules, among the Gemini models in llmwise. They change when the results do.

  • Hard problems: the model that passed the most prompts and, of those, the most hard ones. Models level on both share the pick: the prompts don't tell them apart, so we don't break the tie by price or by name.
  • Best value: among the models that draw on the monthly allowance, the one with the most messages on Pro that passed no more than one prompt fewer than the top model. Ties go to the one that passed more, then to the lower cost per reply.
  • Everyday: among the cheapest models on the page (the everyday models, which come from the daily count, when the page has any), the one that passed the most. Ties go to the one that passed more of the hard prompts, then to the lower cost per reply.

Our math test runs, model by model

How each model did on our 5 math prompts, what each reply counted as on Pro, and what it cost to run.

Our math test runs
ModelPassedHard onesMessages used on ProCost per replyTime per reply
Gemini 3.1 Pro (preview)Google5 of 52 of 21 each, of 125 a month on Pro$0.01007.5 s
Gemini 3.8 FlashGoogle5 of 52 of 21 each, of 250 a month on Pro$0.00143.8 s
Passed: out of the job's 5 prompts, how many replies passed their check (a reply the provider failed to give isn't counted). Messages used: what each reply counted as on Pro, and Pro's count on the model. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Time: from sending to the whole reply.

One prompt, side by side

Each model's reply to “Four-digit numbers whose digits sum to 9”, as it came, with whether it passed and what it cost.

The prompt: Four-digit numbers whose digits sum to 9 (hard)
How many four-digit positive integers have digits that add up to 9?
Show your working briefly, then end with a line that says "Final answer: " followed by the answer alone.
  • Gemini 3.1 Pro

    Passed: Final answer 165: right.

    Let the four-digit positive integer be represented by its digits as $abcd$. We are given that $a + b + c + d = 9$. 
    Since it is a four-digit number, the leading digit $a$ must be at least 1 ($1 \le a \le 9$), while the other digits can be any integer from 0 to 9.
    
    We can use the substitution $x = a - 1$, where $x \ge 0$. 
    Substituting $a = x + 1$ into our equation gives:
    $(x + 1) + b + c + d = 9$
    $x + b + c + d = 8$
    
    We now need to find the number of non-negative integer solutions to this equation. Using the "stars and bars" combinatorial method, the number of solutions for distributing $n$ items among $k$ distinct bins is given by the binomial coefficient $\binom{n + k - 1}{k - 1}$.
    
    Here, $n = 8$ and $k = 4$ (the variables $x, b, c, d$). Since the sum is 8, none of the variables will exceed 9, so no upper limit restrictions are violated.
    Number of solutions = $\binom{8 + 4 - 1}{4 - 1} = \binom{11}{3}$
    
    Calculating $\binom{11}{3}$:
    $\frac{11 \times 10 \times 9}{3 \times 2 \times 1} = \frac{990}{6} = 165$
    
    Final answer: 165

    404 tokens in, 1,181 out (826 of them reasoning) · 9.6 s · $0.0150 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  • Gemini 3.8 Flash

    Passed: Final answer 165: right.

    Let the four digits be represented by $a, b, c,$ and $d$, where $a \in \{1, 2, \dots, 9\}$ and $b, c, d \in \{0, 1, \dots, 9\}$.
    
    We need to find the number of integer solutions to:
    $$a + b + c + d = 9$$
    
    Let $a' = a - 1$ so that $a' \ge 0$. Substituting gives:
    $$a' + b + c + d = 8$$
    
    Since the sum is $8$, no individual variable can exceed $8$, so the constraint that digits must be $\le 9$ is automatically satisfied.
    
    Using the stars and bars formula, the number of non-negative integer solutions is:
    $$\binom{8 + 4 - 1}{4 - 1} = \binom{11}{3} = \frac{11 \times 10 \times 9}{3 \times 2 \times 1} = 165$$
    
    Final answer: 165

    404 tokens in, 495 out (260 of them reasoning) · 4.7 s · $0.0022 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

The math prompts, and how they're scored

Final answer. Automatic. The reply's last “Final answer:” line must hold the right value.

Each prompt was sent the way llmwise sends a message in a side-by-side comparison, which offers no tools: the app's own system prompt, the model's own settings, and Pro's reply size limit (8,000 tokens). Read all five math prompts and how every reply was scored.

Prompts like these to try yourself

Gemini in llmwise

Gemini models in llmwise
ModelOn ProOn FreeContext windowImagesPDFsReasoningAPI price per 1M, in / out
Gemini 3.1 Pro (preview)Google125/mo on ProYes1.05M tokensYesWhole fileYes$2.00 / $12.00
Gemini 3.8 FlashGoogle250/mo on ProYes1.05M tokensYesWhole fileYes$0.75 / $3.75
Each badge is how many messages Pro gets on the model: a month’s, or a day’s on an everyday model. Free is a one-time trial of 5 messages on the models marked. “Text only” models get the text of a PDF, not the file. API prices are the per-token prices in our model catalog as of October 2026 (Google: Google's list price). In llmwise you pay per message, not per token. Gemini 3.1 Pro (preview): the standard rate, for prompts up to 200K tokens. Gemini 3.8 Flash: an introductory price, through December 31, 2026.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

What matters for math

  • Step-by-step reasoning

    Models that reason before answering work through the steps instead of jumping to a number, which matters most on multi-step problems.

  • Checking the arithmetic

    Language models can slip on arithmetic. Running the calculation as code settles it.

  • Showing the work

    A good answer shows its steps clearly enough for you to follow, and flags its assumptions.

Gemini for math in llmwise

  • Readable formulas

    Formulas in answers render as proper math, not raw LaTeX.

  • Run code (paid plans)

    On a paid plan the model can run Python or Node.js in an isolated sandbox once you approve it, read the output and fix what failed. A run counts as 1 Claude Haiku 4.5 message and stops after 60 seconds.

  • Photos of problems

    Attach a photo of a worksheet or a screenshot of a problem to a model that reads images.

  • The Tutor persona

    Pick the Tutor persona to have concepts explained step by step, with checks that you followed.

In llmwise, a message to Gemini goes to its maker, Google, or through OpenRouter when llmwise can't reach the maker directly. See the Privacy Policy.

Tips

  • Ask for the steps, not just the answer.

  • Say what level you're at, so the explanation fits.

  • Ask the model to verify the final answer a second way.

  • For proofs, ask it to list its assumptions.

Questions

Is Gemini good for math?

In our test runs on September 27, 2026, Gemini 3.1 Pro (preview) passed 5 of 5, Gemini 3.8 Flash passed 5 of 5 of our math prompts, against a best result of 5 of 5 among all 19 models. Every reply is published on this page and the methods page, so you can judge them yourself.

Which Gemini model should I use for math?

Start with Gemini 3.8 Flash for coursework, statistics and multi-step word problems, and move up to Gemini 3.1 Pro (preview) for coursework, statistics and multi-step word problems.

Can I use Gemini for math for free?

Yes: Gemini 3.1 Pro (preview) and Gemini 3.8 Flash are on the Free plan (a one-time trial of 5 messages).

Can it check its own answer?

On a paid plan, ask the model to check its answer with code: it runs Python in a sandbox (after you approve) and compares. A run counts as 1 Claude Haiku 4.5 message.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.