Comparison
ChatGPT vs Gemini
In our test runs on September 29, 2026, the same 50 prompts across 10 jobs: GPT's 4 models passed 186 of 200 replies and Gemini's 2 models passed 95 of 100 replies. ChatGPT is OpenAI's own app for its GPT models; llmwise has GPT models in its own chat, not ChatGPT itself.
Based on 300 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
Job by job, both families' best models shared the top result on 8 of the 10 jobs; GPT's alone had it on writing; Gemini's alone had it on customer support.
GPT vs Gemini, job by job
On each job, GPT's pick against Gemini's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.
Writing: GPT-6 Luna passed 5 of 5 and Gemini 3.1 Pro (preview) 4 of 5. GPT-6 Luna answered 3.8× sooner at the median, 2.4 s against 9.1 s. GPT-6 Luna cost 92.3× less, $0.0005 against $0.0459 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 23% longer, in tokens of reply, thinking not counted.
Customer support: GPT-6 Luna passed 4 of 5 and Gemini 3.1 Pro (preview) 5 of 5. GPT-6 Luna answered 6.2× sooner at the median, 1.4 s against 8.7 s. GPT-6 Luna cost 122.5× less, $0.0004 against $0.0529 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 97% longer, in tokens of reply, thinking not counted.
Coding: GPT-6 Luna passed 5 of 5 and Gemini 3.8 Flash 5 of 5. Their median waits were close, 4.6 s against 4.2 s. GPT-6 Luna cost 6.8× less, $0.0014 against $0.0098 for the 5 replies. Gemini 3.8 Flash's replies ran 72% longer, in tokens of reply, thinking not counted.
Math: GPT-6.1 Sol passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6.1 Sol answered 2.7× sooner at the median, 1.8 s against 4.7 s. GPT-6.1 Sol cost 1.7× less, $0.0043 against $0.0071 for the 5 replies. Gemini 3.8 Flash's replies ran 96% longer, in tokens of reply, thinking not counted.
Summarization: GPT-6.1 Sol passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6.1 Sol answered 1.8× sooner at the median, 2.0 s against 3.7 s. Gemini 3.8 Flash cost 1.3× less, $0.0050 against $0.0038 for the 5 replies. Their replies ran to about the same length.
Data analysis: GPT-6 Luna passed 5 of 5 and Gemini 3.1 Pro (preview) 5 of 5. GPT-6 Luna answered 3.5× sooner at the median, 2.5 s against 8.6 s. GPT-6 Luna cost 87.9× less, $0.0009 against $0.0773 for the 5 replies. Gemini 3.1 Pro (preview)'s replies ran 343% longer, in tokens of reply, thinking not counted.
Translation: GPT-6 Luna passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6 Luna answered 1.9× sooner at the median, 1.6 s against 3.0 s. GPT-6 Luna cost 8.1× less, $0.0005 against $0.0037 for the 5 replies. Their replies ran to about the same length.
SQL: GPT-6 Luna passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6 Luna answered 3.1× sooner at the median, 1.2 s against 3.7 s. GPT-6 Luna cost 8.4× less, $0.0004 against $0.0037 for the 5 replies. Gemini 3.8 Flash's replies ran 15% longer, in tokens of reply, thinking not counted.
RAG and answering from documents: GPT-6 Luna passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6 Luna answered 2.5× sooner at the median, 1.3 s against 3.1 s. GPT-6 Luna cost 8.6× less, $0.0004 against $0.0033 for the 5 replies. Gemini 3.8 Flash's replies ran 40% longer, in tokens of reply, thinking not counted.
Agents and tool use: GPT-6 Luna passed 5 of 5 and Gemini 3.8 Flash 5 of 5. GPT-6 Luna answered 2.3× sooner at the median, 1.9 s against 4.4 s. GPT-6 Luna cost 10.7× less, $0.0004 against $0.0043 for the 5 replies. Gemini 3.8 Flash's replies ran 19% longer, in tokens of reply, thinking not counted.
The 2 prompts only one of GPT and Gemini passed
Where one family's pick passed a prompt and the other's didn't, in each check's own words.
Argue both sides of free buses (writing): GPT-6 Luna passed and Gemini 3.1 Pro (preview) didn't. GPT-6 Luna: Graded 4.3 of 5 on average (lowest 4). Gemini 3.1 Pro (preview): Graded 3.7 of 5 on average (lowest 3).
A frustrated customer (customer support): Gemini 3.1 Pro (preview) passed and GPT-6 Luna didn't. GPT-6 Luna: Graded 3.0 of 5 on average (lowest 2). Gemini 3.1 Pro (preview): Graded 4.0 of 5 on average (lowest 3).
Each GPT model against each Gemini model
Every GPT model against every Gemini model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.
GPT-6 Astra vs Gemini 3.1 Pro: 48 and 49 of 50; 3 prompts split them; their replies cost about the same in all, and GPT-6 Astra answered sooner on 50, Gemini 3.1 Pro on 0.
GPT-6 Astra vs Gemini 3.8 Flash: 48 and 46 of 50; 4 prompts split them; Gemini 3.8 Flash's replies cost 9.1× less in all, and GPT-6 Astra answered sooner on 32, Gemini 3.8 Flash on 14.
GPT-6.1 Sol vs Gemini 3.1 Pro: 46 and 49 of 50; 5 prompts split them; GPT-6.1 Sol's replies cost 8.9× less in all, and GPT-6.1 Sol answered sooner on 50, Gemini 3.1 Pro on 0.
GPT-6.1 Sol vs Gemini 3.8 Flash: 46 and 46 of 50; 4 prompts split them; their replies cost about the same in all, and GPT-6.1 Sol answered sooner on 40, Gemini 3.8 Flash on 7.
GPT-6 Sol vs Gemini 3.1 Pro: 45 and 49 of 50; 6 prompts split them; GPT-6 Sol's replies cost 4.1× less in all, and GPT-6 Sol answered sooner on 50, Gemini 3.1 Pro on 0. GPT-6 Sol vs Gemini 3.1 Pro (preview).
GPT-6 Sol vs Gemini 3.8 Flash: 45 and 46 of 50; 5 prompts split them; Gemini 3.8 Flash's replies cost 2.0× less in all, and GPT-6 Sol answered sooner on 37, Gemini 3.8 Flash on 9. GPT-6 Sol vs Gemini 3.8 Flash.
GPT-6 Luna vs Gemini 3.1 Pro: 47 and 49 of 50; 4 prompts split them; GPT-6 Luna's replies cost 90.4× less in all, and GPT-6 Luna answered sooner on 50, Gemini 3.1 Pro on 0.
GPT-6 Luna vs Gemini 3.8 Flash: 47 and 46 of 50; 5 prompts split them; GPT-6 Luna's replies cost 10.9× less in all, and GPT-6 Luna answered sooner on 43, Gemini 3.8 Flash on 6. GPT-6 Luna vs Gemini 3.8 Flash.
One writing prompt, both replies
GPT-6 Luna and Gemini 3.1 Pro (preview) on a writing prompt, the job where they differed most: both replies as they came.
A hard prompt: “Argue both sides of free buses”, in full.
GPT-6 Luna
Passed: Graded 4.3 of 5 on average (lowest 4).
Making buses free would give every resident reliable access to jobs, school, health care, and daily necessities, regardless of income. It would also speed boarding, reduce fare-enforcement costs, and encourage people to leave cars at home—easing congestion, pollution, and parking pressure. A city that treats public transit as essential infrastructure, rather than a service people must qualify to afford, can build a more connected and equitable community. But “free” buses still have to be paid for, and the bill could divert money from more urgent improvements: frequent service, safer stops, cleaner vehicles, and routes to underserved neighborhoods. If buses remain slow, crowded, or unreliable, removing fares may do little to persuade drivers to switch, while reducing funds that could make transit genuinely useful. Targeted fare assistance can protect riders who need it most without subsidizing every trip, including those by people who can readily afford to pay.
384 tokens in, 186 out · 2.5 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·
Gemini 3.1 Pro
Failed: Graded 3.7 of 5 on average (lowest 3).
Making city buses free is a powerful catalyst for economic equality and environmental sustainability. By eliminating fares, cities immediately remove a significant financial burden from low-income residents, improving their access to jobs, education, and healthcare. Furthermore, free transit incentivizes people to leave their cars at home, dramatically reducing traffic congestion and cutting harmful greenhouse gas emissions. This shift not only improves urban air quality but also reduces the long-term wear and tear on city infrastructure, ultimately creating a cleaner, fairer, and more accessible environment for all residents. However, eliminating fares fundamentally undermines the financial stability and quality of the transit system itself. Fares provide crucial revenue required for essential maintenance, route expansion, and driver salaries. Without this income, cities are forced to either slash service frequencies and coverage or drastically raise local taxes to cover the deficit. Consequently, free transit often results in overcrowded, unreliable, and dilapidated buses. Passengers ultimately prioritize speed, safety, and reliability over zero cost; a free bus is useless to a commuter if it never arrives on time.
409 tokens in, 699 out (492 of them reasoning) · 9.1 s · $0.0092 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·
Every model, every job
All 6 GPT and Gemini models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.
| Model | Passed | Hard ones | Cost per reply | On Pro |
|---|---|---|---|---|
| GPT-6 AstraOpenAI | 48 of 50 | 20 of 20 | $0.0117 | Up to 31 a month |
| GPT-6.1 SolOpenAI | 46 of 50 | 19 of 20 | $0.0012 | Up to 125 a month |
| GPT-6 SolOpenAI | 45 of 50 | 19 of 20 | $0.0026 | Up to 125 a month |
| GPT-6 LunaOpenAI | 47 of 50 | 20 of 20 | $0.00012 | Up to 60 a day |
| Gemini 3.1 Pro (preview)Google | 49 of 50 | 19 of 20 | $0.0107 | Up to 125 a month |
| Gemini 3.8 FlashGoogle | 46 of 50 | 18 of 20 | $0.0013 | Up to 250 a month |
Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
Their own subscriptions
Each company's own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.
ChatGPT Plus (OpenAI), $20 a month
Models it names: GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. On its limits: “To ensure a smooth experience for all users, Plus subscriptions may include usage limits such as message caps, especially during high demand. These limits may vary based on system conditions.” ChatGPT Plus vs llmwise.
Checked : OpenAI Help Center: What is ChatGPT Plus?, OpenAI Help Center: Managing usage with GPT-6 Astra in Work and Codex and ChatGPT pricing.
Google AI Pro (Google), $19.99 a month
Models it names: Gemini 3.1 Pro. On its limits: “AI Pro: 4x higher than standard limits” Google AI Pro vs llmwise.
Checked : Google One: Google AI plans and Gemini Apps Help: limits and upgrades for Google AI subscribers.
llmwise Pro, $20 a month, has all 6 of these models in one chat, on one monthly allowance. On it: GPT-6.1 Sol up to 125 messages a month and Gemini 3.1 Pro (preview) up to 125.
The lineups at a glance
What follows from each model's facts in our catalog.
The lineups
GPT: 4 models, GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna. Gemini: 2 models, Gemini 3.1 Pro (preview) and Gemini 3.8 Flash.
Price per message
The least expensive GPT model is GPT-6 Luna (60 messages a day on Pro); the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro).
Context window
GPT goes up to 1.05M tokens (GPT-6 Astra); Gemini up to 1.05M tokens (Gemini 3.1 Pro (preview)).
Images and PDFs
Every model here reads images. Every model here takes a PDF as a whole file.
On the Free plan
Free's one-time trial of 5 messages covers GPT-6.1 Sol, GPT-6 Sol, GPT-6 Luna, Gemini 3.1 Pro (preview), and Gemini 3.8 Flash. Paid plans have every model, with messages every month.
Model by model
Two named models side by side, prompt by prompt, each with its messages on every plan.
More head-to-heads
Each of GPT and Gemini against the other families, every job from the same test runs.
Where your messages go
In llmwise, a message to GPT or Gemini goes to the model's maker, or through OpenRouter when llmwise can't reach the maker directly. The Privacy Policy has the details.
Questions
Which is better, ChatGPT or Gemini?
In our test runs on September 29, 2026, the same 50 prompts across 10 jobs: GPT's 4 models passed 186 of 200 replies and Gemini's 2 models passed 95 of 100 replies. Job by job, both families' best models shared the top result on 8 of the 10 jobs; GPT's alone had it on writing; Gemini's alone had it on customer support.
Which is cheaper, GPT or Gemini?
In llmwise, the least expensive GPT model is GPT-6 Luna (60 messages a day on Pro), and the least expensive Gemini model is Gemini 3.8 Flash (250 messages a month on Pro). At API list prices (October 2026), a typical message of 4,000 tokens in and 700 out costs $0.0008 on GPT-6 Luna and $0.0056 on Gemini 3.8 Flash.
Can I use GPT and Gemini in the same chat?
Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.