Compare answers
Compare AI answers side by side
Compare sends one question to 2 or 3 models at once; each answer counts as one message of its model, and picking one to go on with is free. On one everyday customer-service prompt in our tests, 8 of 15 models gave a reply our grader failed.
A real answer from our test runs
GPT-6 LunaYou roll three fair six-sided dice. What's the probability that the highest number showing is exactly 5? Give it as a fraction in lowest terms. Show your working briefly, then end with a line that says "Final answer: " followed by the answer alone.
GPT-6 Luna · passed our check · September 27, 2026
There are \(5^3\) outcomes where every die is at most 5, and \(4^3\) where every die is at most 4. Thus the probability is
\[
\frac{5^3-4^3}{6^3}=\frac{61}{216}.
\]
Final answer: 61/216Ask your own question: sign up free for 5 messages on 12 of the models, no card.
Start freeTest runs and prices checked . Updated .
Short answer
Sign in, choose Compare models, and send one question to 2 or 3 models. It counts as one message of each model, said before you send; going on with one answer is free.
How Compare works
A signed-in feature on every plan; on the free trial, each answer is one of your free messages.
Sign in with an email link and choose Compare models in the sidebar: a new chat opens with Compare on.
Pick 2 or 3 models. The message box shows what the compare counts as, model by model.
Send the question once. Compare sends a plain question: no tools, web search or images.
Read the answers side by side and pick one to go on with. Picking an answer to go on with is free.
One question, three answers
The same everyday customer support prompt, “A frustrated customer”, sent to GPT-6 Luna, Claude Sonnet 5, and Gemini 3.1 Pro: GPT-6 Luna failed, Claude Sonnet 5 passed, and Gemini 3.1 Pro passed.
The prompt: A frustrated customer (everyday)
Northwind Outfitters support policy - Returns: within 30 days of delivery, unworn items with tags get a full refund to the original payment method. - After 30 days and up to 60 days: exchange or a store gift card only, no refund. - Refunds are paid within 5 business days of the return reaching our warehouse. - Standard delivery takes 3 to 5 business days. If an order is more than 7 business days late, the customer gets free express shipping on their next order. - Staff can't give discount codes worth more than 15%. Customer message: "This is the THIRD time you've sent the wrong size. I ordered a medium and got a large AGAIN. I'm so fed up with this." Write the reply to the customer, following the policy. At most 150 words, signed "Sam, Northwind support".
GPT-6 Luna
Failed: Graded 3.0 of 5 on average (lowest 2).
Hi, I’m really sorry we’ve sent the wrong size again. I understand how frustrating this is, especially after three mistakes. We’ll help make it right—please reply with your order number so we can arrange the appropriate return or exchange under our policy. Sam, Northwind support
516 tokens in, 61 out · 1.4 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·
Claude Sonnet 5
Passed: Graded 4.3 of 5 on average (lowest 4).
Hi there, I'm really sorry — three wrong-size shipments in a row is completely unacceptable, and I understand your frustration. Let's fix this properly this time. Please send me your order number, and I'll personally double-check the medium is correctly noted before it ships. For the incorrect large you received, you have two options: 1. **Return it for a full refund** (since it's unworn with tags, this qualifies for our standard return policy), refunded within 5 business days of it reaching our warehouse, or 2. **Exchange it for the correct medium** at no extra cost, with a prepaid return label sent to you now. Given the repeated mix-ups, I'd also like to offer you a 15% discount code on your next order as an apology. Please let me know which option works best, and I'll get it sorted right away. Sam, Northwind support
718 tokens in, 270 out · 4.4 s · $0.0041 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·
Gemini 3.1 Pro
Passed: Graded 4.0 of 5 on average (lowest 3).
Hi there, I am so sorry. You have every right to be frustrated. Sending you the wrong size three times is completely unacceptable, and I sincerely apologize for the repeated hassle. I want to make this right immediately. We can process a free exchange to finally get you the correct Medium. Alternatively, if you prefer to return the unworn item with its tags within 30 days, we will issue a full refund to your original payment method (processed within 5 business days of it reaching our warehouse). Please let me know which option you prefer, and I will send you a prepaid return shipping label right away. To help make up for this frustrating experience, please use the code FIXIT15 for 15% off your next order. Again, I am so sorry for the inconvenience. Best, Sam, Northwind support
557 tokens in, 821 out (643 of them reasoning) · 9.8 s · $0.0110 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·
Where our models disagreed
Every prompt of our 50 that at least one of the 15 models failed, and whose replies failed it.
| Prompt | Job | Failed | Models that failed it |
|---|---|---|---|
| A frustrated customerEveryday | Customer support | 8 of 15 | Claude Haiku 4.5, GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, Gemini 3.8 Flash, DeepSeek V4 Pro, Grok 4.7, and GLM 5.3 Flash |
| An article in three bulletsEveryday | Summarization | 7 of 15 | Claude Fable 5.1, Claude Sonnet 5, GPT-6 Sol, GPT-6 Luna, Kimi K3, GLM 5.3, and GLM 5.3 Flash |
| A product announcement with five rulesHard | Writing | 6 of 15 | Claude Fable 5.1, Claude Haiku 4.5, Gemini 3.8 Flash, DeepSeek V4 Pro, Kimi K3, and GLM 5.3 Flash |
| Argue both sides of free busesHard | Writing | 6 of 15 | Claude Fable 5.1, Claude Opus 5.5, Gemini 3.1 Pro, DeepSeek V4 Pro, GLM 5.3, and GLM 5.3 Flash |
| Rewrite corporate jargon in plain wordsEveryday | Writing | 5 of 15 | GPT-6 Sol, Gemini 3.8 Flash, DeepSeek V4.1 Flash, Grok 4.7, and GLM 5.3 Flash |
| An email thread in one sentenceEveryday | Summarization | 5 of 15 | Claude Fable 5.1, Claude Sonnet 5, DeepSeek V4 Pro, Kimi K3, and GLM 5.3 Flash |
| Correlation between ad spend and sign-upsHard | Data analysis | 4 of 15 | Claude Sonnet 5, Claude Haiku 4.5, Gemini 3.8 Flash, and GLM 5.3 Flash |
| A late orderEveryday | Customer support | 3 of 15 | GPT-6 Astra, GPT-6 Sol, and GLM 5.3 Flash |
| Parse CSV with quoted fieldsHard | Coding | 2 of 15 | Claude Haiku 4.5 and DeepSeek V4 Pro |
| Announce a second bakery shop on LinkedInEveryday | Writing | 2 of 15 | Claude Sonnet 5 and Grok 4.7 |
| A refund request outside the windowHard | Customer support | 2 of 15 | Kimi K3 and GLM 5.3 Flash |
| A message with a planted instructionHard | Customer support | 2 of 15 | Claude Fable 5.1 and Claude Haiku 4.5 |
| Evaluate an arithmetic expression, no evalHard | Coding | 1 of 15 | Claude Haiku 4.5 |
| Pens at 3 for $4Everyday | Math | 1 of 15 | GPT-6 Luna |
| Average order value in AugustEveryday | Data analysis | 1 of 15 | Claude Haiku 4.5 |
| Idioms into natural JapaneseHard | Translation | 1 of 15 | GPT-6 Sol |
| Book a meeting from a sentenceEveryday | Agents and tool use | 1 of 15 | GLM 5.3 Flash |
More ways to chat
The other chat pages, each with its own models and test results.
- AI chat with every top model in one place
- Free AI chat, and exactly what's free
- Ask AI anything, and ask more than one
- An AI assistant with every top model, and you pick which
- Talk to AI, in a conversation that keeps its thread
- Free GPT-6 chat: try GPT-6 Luna and Sol
- Multi-model AI chat: every family, one subscription
- Chat with Claude online
- Chat with GPT-6 online
- Chat with Gemini online
- Chat with Grok online
- Chat with DeepSeek online
- Chat with Kimi online
Questions
How do I compare AI answers side by side?
Sign in, choose Compare models in the sidebar, pick 2 or 3 models and send your question. The answers arrive side by side; pick the one to go on with, and the chat continues on that model.
What does a compare cost?
Exactly the sum of each model's own message: a compare of Claude Sonnet 5, GPT-6 Sol and Gemini 3.1 Pro counts as one message of each, nothing more. The message box says the total before you send. Picking an answer to go on with is free.
Can I compare on the free trial?
Yes. On the Free trial each model's answer is one of your 5 free messages, so a compare of 3 uses 3 of them; the box shows how many you have left before you send.
Can a compare search the web or make images?
No: a compare sends a plain question, with no tools, web search or images, so what it counts as is exactly what the box says.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.