Skip to content

Comparison

ChatGPT vs Gemini

In our test runs on September 29, 2026, the same 50 prompts across 10 jobs: GPT's 4 models passed 186 of 200 replies and Gemini's 2 models passed 95 of 100 replies. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself.

Based on 300 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

Job by job, both families' best models shared the top result on 8 of the 10 jobs; GPT's alone had it on writing; Gemini's alone had it on customer support.

GPT vs Gemini, job by job

On each job, GPT's pick against Gemini's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.

  • Writing: GPT-6 Luna passed 5 of 5 and Gemini 3.1 Pro (preview) 4 of 5. GPT-6 Luna answered 3.8× sooner at the median, 2.4 s against 9.1 s. GPT-6 Luna cost 92.3× less, $0.0005 against $0.0459 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 23% longer, in tokens of reply, thinking not counted.

  • Customer support: GPT-6 Luna passed 4 of 5 and Gemini 3.1 Pro (preview) 5 of 5. GPT-6 Luna answered 6.2× sooner at the median, 1.4 s against 8.7 s. GPT-6 Luna cost 122.5× less, $0.0004 against $0.0529 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 97% longer, in tokens of reply, thinking not counted.

  • Coding: GPT-6 Luna passed 5 of 5 and Gemini 3.8 Flash 5 of 5. Their median waits were close, 4.6 s against 4.2 s. GPT-6 Luna cost 6.8× less, $0.0014 against $0.0098 for the 5 replies. Gemini 3.8 Flash's replies ran 72% longer, in tokens of reply, thinking not counted.

  • Math: GPT-6.1 Sol passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6.1 Sol answered 2.7× sooner at the median, 1.8 s against 4.7 s. GPT-6.1 Sol cost 1.7× less, $0.0043 against $0.0071 for the 5 replies. Gemini 3.8 Flash's replies ran 96% longer, in tokens of reply, thinking not counted.

  • Summarization: GPT-6.1 Sol passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6.1 Sol answered 1.8× sooner at the median, 2.0 s against 3.7 s. Gemini 3.8 Flash cost 1.3× less, $0.0050 against $0.0038 for the 5 replies. Their replies ran to about the same length.

  • Data analysis: GPT-6 Luna passed 5 of 5 and Gemini 3.1 Pro (preview) 5 of 5. GPT-6 Luna answered 3.5× sooner at the median, 2.5 s against 8.6 s. GPT-6 Luna cost 87.9× less, $0.0009 against $0.0773 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 343% longer, in tokens of reply, thinking not counted.

  • Translation: GPT-6 Luna passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6 Luna answered 1.9× sooner at the median, 1.6 s against 3.0 s. GPT-6 Luna cost 8.1× less, $0.0005 against $0.0037 for the 5 replies. Their replies ran to about the same length.

  • SQL: GPT-6 Luna passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6 Luna answered 3.1× sooner at the median, 1.2 s against 3.7 s. GPT-6 Luna cost 8.4× less, $0.0004 against $0.0037 for the 5 replies. Gemini 3.8 Flash's replies ran 15% longer, in tokens of reply, thinking not counted.

  • RAG and answering from documents: GPT-6 Luna passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6 Luna answered 2.5× sooner at the median, 1.3 s against 3.1 s. GPT-6 Luna cost 8.6× less, $0.0004 against $0.0033 for the 5 replies. Gemini 3.8 Flash's replies ran 40% longer, in tokens of reply, thinking not counted.

  • Agents and tool use: GPT-6 Luna passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6 Luna answered 2.3× sooner at the median, 1.9 s against 4.4 s. GPT-6 Luna cost 10.7× less, $0.0004 against $0.0043 for the 5 replies. Gemini 3.8 Flash's replies ran 19% longer, in tokens of reply, thinking not counted.

The 2 prompts only one of GPT and Gemini passed

Where one family's pick passed a prompt and the other's didn't, in each check's own words.

  • Argue both sides of free buses (writing): GPT-6 Luna passed and Gemini 3.1 Pro (preview) didn't. GPT-6 Luna: Graded 4.3 of 5 on average (lowest 4). Gemini 3.1 Pro (preview): Graded 3.7 of 5 on average (lowest 3).

  • A frustrated customer (customer support): Gemini 3.1 Pro (preview) passed and GPT-6 Luna didn't. GPT-6 Luna: Graded 3.0 of 5 on average (lowest 2). Gemini 3.1 Pro (preview): Graded 4.0 of 5 on average (lowest 3).

Each GPT model against each Gemini model

Every GPT model against every Gemini model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.

  • GPT-6 Astra vs Gemini 3.1 Pro: 48 and 49 of 50; 3 prompts split them; their replies cost about the same in all, and GPT-6 Astra answered sooner on 50, Gemini 3.1 Pro on 0.

  • GPT-6 Astra vs Gemini 3.8 Flash: 48 and 46 of 50; 4 prompts split them; Gemini 3.8 Flash's replies cost 9.1× less in all, and GPT-6 Astra answered sooner on 32, Gemini 3.8 Flash on 14.

  • GPT-6.1 Sol vs Gemini 3.1 Pro: 46 and 49 of 50; 5 prompts split them; GPT-6.1 Sol's replies cost 8.9× less in all, and GPT-6.1 Sol answered sooner on 50, Gemini 3.1 Pro on 0.

  • GPT-6.1 Sol vs Gemini 3.8 Flash: 46 and 46 of 50; 4 prompts split them; their replies cost about the same in all, and GPT-6.1 Sol answered sooner on 40, Gemini 3.8 Flash on 7.

  • GPT-6 Sol vs Gemini 3.1 Pro: 45 and 49 of 50; 6 prompts split them; GPT-6 Sol's replies cost 4.1× less in all, and GPT-6 Sol answered sooner on 50, Gemini 3.1 Pro on 0. GPT-6 Sol vs Gemini 3.1 Pro (preview).

  • GPT-6 Sol vs Gemini 3.8 Flash: 45 and 46 of 50; 5 prompts split them; Gemini 3.8 Flash's replies cost 2.0× less in all, and GPT-6 Sol answered sooner on 37, Gemini 3.8 Flash on 9. GPT-6 Sol vs Gemini 3.8 Flash.

  • GPT-6 Luna vs Gemini 3.1 Pro: 47 and 49 of 50; 4 prompts split them; GPT-6 Luna's replies cost 90.4× less in all, and GPT-6 Luna answered sooner on 50, Gemini 3.1 Pro on 0.

  • GPT-6 Luna vs Gemini 3.8 Flash: 47 and 46 of 50; 5 prompts split them; GPT-6 Luna's replies cost 10.9× less in all, and GPT-6 Luna answered sooner on 43, Gemini 3.8 Flash on 6. GPT-6 Luna vs Gemini 3.8 Flash.

One writing prompt, both replies

GPT-6 Luna and Gemini 3.1 Pro (preview) on a writing prompt, the job where they differed most: both replies as they came.

A hard prompt: “Argue both sides of free buses”, in full.

  • GPT-6 Luna

    Passed: Graded 4.3 of 5 on average (lowest 4).

    Making buses free would give every resident reliable access to jobs, school, health care, and daily necessities, regardless of income. It would also speed boarding, reduce fare-enforcement costs, and encourage people to leave cars at home—easing congestion, pollution, and parking pressure. A city that treats public transit as essential infrastructure, rather than a service people must qualify to afford, can build a more connected and equitable community.
    
    But “free” buses still have to be paid for, and the bill could divert money from more urgent improvements: frequent service, safer stops, cleaner vehicles, and routes to underserved neighborhoods. If buses remain slow, crowded, or unreliable, removing fares may do little to persuade drivers to switch, while reducing funds that could make transit genuinely useful. Targeted fare assistance can protect riders who need it most without subsidizing every trip, including those by people who can readily afford to pay.

    384 tokens in, 186 out · 2.5 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  • Gemini 3.1 Pro

    Failed: Graded 3.7 of 5 on average (lowest 3).

    Making city buses free is a powerful catalyst for economic equality and environmental sustainability. By eliminating fares, cities immediately remove a significant financial burden from low-income residents, improving their access to jobs, education, and healthcare. Furthermore, free transit incentivizes people to leave their cars at home, dramatically reducing traffic congestion and cutting harmful greenhouse gas emissions. This shift not only improves urban air quality but also reduces the long-term wear and tear on city infrastructure, ultimately creating a cleaner, fairer, and more accessible environment for all residents.
    
    However, eliminating fares fundamentally undermines the financial stability and quality of the transit system itself. Fares provide crucial revenue required for essential maintenance, route expansion, and driver salaries. Without this income, cities are forced to either slash service frequencies and coverage or drastically raise local taxes to cover the deficit. Consequently, free transit often results in overcrowded, unreliable, and dilapidated buses. Passengers ultimately prioritize speed, safety, and reliability over zero cost; a free bus is useless to a commuter if it never arrives on time.

    409 tokens in, 699 out (492 of them reasoning) · 9.1 s · $0.0092 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

Every model, every job

All 6 GPT and Gemini models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.

Every GPT and Gemini model across our test runs
ModelPassedHard onesCost per replyOn Pro
GPT-6 AstraOpenAI48 of 5020 of 20$0.0117Up to 31 a month
GPT-6.1 SolOpenAI46 of 5019 of 20$0.0012Up to 125 a month
GPT-6 SolOpenAI45 of 5019 of 20$0.0026Up to 125 a month
GPT-6 LunaOpenAI47 of 5020 of 20$0.00012Up to 60 a day
Gemini 3.1 Pro (preview)Google49 of 5019 of 20$0.0107Up to 125 a month
Gemini 3.8 FlashGoogle46 of 5018 of 20$0.0013Up to 250 a month

Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Their own subscriptions

Each company's own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.

llmwise Pro, $20 a month, has all 6 of these models in one chat, on one monthly allowance. On it: GPT-6.1 Sol up to 125 messages a month and Gemini 3.1 Pro (preview) up to 125.

The lineups at a glance

What follows from each model's facts in our catalog.

  • The lineups

    GPT: 4 models, GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna. Gemini: 2 models, Gemini 3.1 Pro (preview) and Gemini 3.8 Flash.

  • Price per message

    The least expensive GPT model is GPT-6 Luna (60 messages a day on Pro); the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro).

  • Context window

    GPT goes up to 1.05M tokens (GPT-6 Astra); Gemini up to 1.05M tokens (Gemini 3.1 Pro (preview)).

  • Images and PDFs

    Every model here reads images. Every model here takes a PDF as a whole file.

  • On the Free plan

    Free's one-time trial of 5 messages covers GPT-6.1 Sol, GPT-6 Sol, GPT-6 Luna, Gemini 3.1 Pro (preview), and Gemini 3.8 Flash. Paid plans have every model, with messages every month.

Model by model

Two named models side by side, prompt by prompt, each with its messages on every plan.

More head-to-heads

Each of GPT and Gemini against the other families, every job from the same test runs.

Where your messages go

In llmwise, a message to GPT or Gemini goes to the model's maker, or through OpenRouter when llmwise can't reach the maker directly. The Privacy Policy has the details.

Bar chart: Prompts passed in our test runs, all 10 jobs. Gemini 3.1 Pro (preview): 49 of 50; GPT-6 Astra: 48 of 50; GPT-6 Luna: 47 of 50; GPT-6.1 Sol: 46 of 50; Gemini 3.8 Flash: 46 of 50; GPT-6 Sol: 45 of 50.
Our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026: the same prompts for every model, each reply checked the same way.

Questions

Which is better, ChatGPT or Gemini?

In our test runs on September 29, 2026, the same 50 prompts across 10 jobs: GPT's 4 models passed 186 of 200 replies and Gemini's 2 models passed 95 of 100 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; GPT's alone had it on writing; Gemini's alone had it on customer support.

Which is cheaper, GPT or Gemini?

In llmwise, the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro), and the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro). At API list prices (October 2026), a typical message of 4,000 tokens in and 700 out costs $0.0008 on GPT-6 Luna and $0.0056 on Gemini 3.8 Flash.

Can I use GPT and Gemini in the same chat?

Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.