Tested prompt · Summarization
An email thread in one sentence: every AI model's reply, tested
We sent this everyday summarization prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.
Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
10 of 16 models passed this summarization prompt's check (rubric); Claude Fable 5.1, Claude Sonnet 5.5, Claude Sonnet 5, DeepSeek V4 Pro, Kimi K3, and GLM 5.3 Flash didn't. The cheapest reply that passed was GPT-6 Luna's, at $0.000067; the fastest, GLM 5.3's in 1.0 s. The dearest reply, Claude Fable 5.1's, cost 170 times as much ($0.0114).
The prompt, as sent, and its check
Checked by rubric (graded), the same way for every model.
An email thread in one sentence (everyday)
Give a one-sentence summary, at most 30 words, of this email thread.
From Rosa (Mon 9:02): The printer says the brochures can't ship until the 18th, not the 11th. Paper shortage.
From Idris (Mon 9:40): The trade fair is on the 20th, so the 18th still works if we pick them up ourselves.
From Rosa (Mon 10:15): Fine by me. I'll drive to the printer on the 18th. Can someone book the van?
From Idris (Mon 10:31): Booked the van for the 18th, 8:00 to 12:00.
Graded on:
- Gets the outcome: Says the brochures are late but will be picked up on the 18th, in time for the fair.
- Accurate: Nothing wrong or invented.
- Clear: One clear sentence.
Exactly what this prompt's replies are checked against, with every other prompt of our test runs.
Every model's result
All 16 models on this prompt, in catalog order.
| Model | Result | Cost | Time | Reply |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | Failed: Graded 5.0 of 5 on average (lowest 5); but 33 words, over the 30 allowed. | $0.0114 | 2.5 s | 67 tokens |
| Claude Opus 5.5Anthropic | Passed: Graded 5.0 of 5 on average (lowest 5). | $0.0044 | 2.6 s | 61 tokens |
| Claude Sonnet 5.5Anthropic | Failed: Graded 4.7 of 5 on average (lowest 4); but 35 words, over the 30 allowed. | $0.0023 | 1.4 s | 69 tokens |
| Claude Sonnet 5Anthropic | Failed: Graded 5.0 of 5 on average (lowest 5); but 32 words, over the 30 allowed. | $0.0019 | 2.3 s | 61 tokens |
| Claude Haiku 4.5Anthropic | Passed: Graded 4.7 of 5 on average (lowest 4). | $0.00072 | 1.3 s | 41 tokens |
| GPT-6 AstraOpenAI | Passed: Graded 5.0 of 5 on average (lowest 5). | $0.0071 | 1.9 s | 45 tokens |
| GPT-6 SolOpenAI | Passed: Graded 5.0 of 5 on average (lowest 5). | $0.0014 | 2.0 s | 39 tokens |
| GPT-6 LunaOpenAI | Passed: Graded 4.3 of 5 on average (lowest 3). | $0.000067 | 1.1 s | 37 tokens |
| Gemini 3.1 Pro (preview)Google | Passed: Graded 4.7 of 5 on average (lowest 4). | $0.0058 | 6.3 s | 25 tokens |
| Gemini 3.8 FlashGoogle | Passed: Graded 4.7 of 5 on average (lowest 4). | $0.00053 | 3.8 s | 35 tokens |
| DeepSeek V4.1 FlashDeepSeek | Passed: Graded 4.7 of 5 on average (lowest 4). | $0.00011 | 1.4 s | 38 tokens |
| DeepSeek V4 ProDeepSeek | Failed: Graded 4.7 of 5 on average (lowest 4); but 33 words, over the 30 allowed. | $0.00015 | 2.8 s | 48 tokens |
| Grok 4.7xAI | Passed: Graded 4.3 of 5 on average (lowest 4). | $0.0018 | 1.7 s | 34 tokens |
| Kimi K3Moonshot | Failed: Graded 5.0 of 5 on average (lowest 5); but 32 words, over the 30 allowed. | $0.0030 | 6.4 s | 56 tokens |
| GLM 5.3Z.ai | Passed: Graded 4.3 of 5 on average (lowest 3). | $0.00023 | 1.0 s | 37 tokens |
| GLM 5.3 FlashZ.ai | Failed: Graded 3.7 of 5 on average (lowest 3). | $0.00014 | 4.0 s | 46 tokens |
Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.
Every reply
The 6 replies that failed first, then the ones that passed, from the cheapest up.
GLM 5.3 Flash
Failed: Graded 3.7 of 5 on average (lowest 3).
Brochure delivery was delayed to the 18th due to a paper shortage, but the team will self-pick them up that morning, with Idris booking a van for 8:00–12:00.
504 tokens in, 58 out (12 of them reasoning) · 4.0 s · $0.0001 · 1 message on Pro · answered by z-ai/glm-5.3-flash via Wafer ·
DeepSeek V4 Pro
Failed: Graded 4.7 of 5 on average (lowest 4); but 33 words, over the 30 allowed.
The brochures won't ship until the 18th due to a paper shortage, but that still works for the 20th trade fair; Rosa will pick them up and Idris booked the van for 8–12.
497 tokens in, 117 out (69 of them reasoning) · 2.8 s · $0.0001 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Baidu ·
Claude Sonnet 5
Failed: Graded 5.0 of 5 on average (lowest 5); but 32 words, over the 30 allowed.
Brochure delivery is delayed to the 18th due to a paper shortage, but Rosa will pick them up herself using a van Idris booked, still in time for the 20th trade fair.
652 tokens in, 61 out · 2.3 s · $0.0019 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·
Claude Sonnet 5.5
Failed: Graded 4.7 of 5 on average (lowest 4); but 35 words, over the 30 allowed.
Due to a paper shortage, the brochures won't ship until the 18th, so Rosa will pick them up that day in a van Idris booked (8:00–12:00), in time for the trade fair on the 20th.
656 tokens in, 69 out · 1.4 s · $0.0023 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·
Kimi K3
Failed: Graded 5.0 of 5 on average (lowest 5); but 32 words, over the 30 allowed.
Due to a paper shortage delaying brochure shipping to the 18th, Rosa and Idris arranged to pick them up themselves, booking a van for the 18th ahead of the 20th trade fair.
574 tokens in, 65 out (9 of them reasoning) · 6.4 s · $0.0030 · 1 message on Pro · answered by moonshotai/kimi-k3 via Fireworks ·
Claude Fable 5.1
Failed: Graded 5.0 of 5 on average (lowest 5); but 33 words, over the 30 allowed.
Brochure shipping slipped from the 11th to the 18th due to a paper shortage, so Rosa will collect them on the 18th in a van Idris booked, ahead of the 20th trade fair.
654 tokens in, 67 out · 2.5 s · $0.0114 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·
GPT-6 Luna
Passed: Graded 4.3 of 5 on average (lowest 3).
The brochures will be ready on the 18th, and Rosa will pick them up using the van Idris booked for 8:00–12:00.
483 tokens in, 37 out · 1.1 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·
DeepSeek V4.1 Flash
Passed: Graded 4.7 of 5 on average (lowest 4).
Brochure printing was delayed to the 18th by a paper shortage, but Rosa will pick them up that day for the 20th trade fair, with the van booked.
525 tokens in, 41 out (3 of them reasoning) · 1.4 s · $0.0001 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via StreamLake ·
GLM 5.3
Passed: Graded 4.3 of 5 on average (lowest 3).
The brochure shipment was delayed to the 18th due to a paper shortage, so Rosa will pick them up herself, with Idris booking a van for that morning.
503 tokens in, 37 out · 1.0 s · $0.0002 · 1 message on Pro · answered by z-ai/glm-5.3 via Baidu ·
Gemini 3.8 Flash
Passed: Graded 4.7 of 5 on average (lowest 4).
After brochure shipping was delayed to the 18th due to a paper shortage, Rosa and Idris arranged to pick them up themselves using a booked van before the trade fair.
529 tokens in, 35 out · 3.8 s · $0.0005 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·
Claude Haiku 4.5
Passed: Graded 4.7 of 5 on average (lowest 4).
Printer delayed brochures to the 18th due to paper shortage; Rosa will pick them up with a booked van, arriving in time for the 20th trade fair.
520 tokens in, 41 out · 1.3 s · $0.0007 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·
GPT-6 Sol
Passed: Graded 5.0 of 5 on average (lowest 5).
A paper shortage delayed the brochures until the 18th, but Rosa will collect them in the van Idris booked, ahead of the trade fair on the 20th.
483 tokens in, 39 out · 2.0 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Grok 4.7
Passed: Graded 4.3 of 5 on average (lowest 4).
Brochures ship the 18th due to a paper shortage; Idris booked a van that morning so Rosa can pick them up before the 20th fair.
1,720 tokens in, 90 out (56 of them reasoning) · 1.7 s · $0.0018 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
Claude Opus 5.5
Passed: Graded 5.0 of 5 on average (lowest 5).
Due to a paper shortage, brochures ship on the 18th instead of the 11th, so Rosa will collect them using a van Idris booked, ahead of the 20th trade fair.
654 tokens in, 61 out · 2.6 s · $0.0044 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·
Gemini 3.1 Pro
Passed: Graded 4.7 of 5 on average (lowest 4).
After a printing delay, Idris booked a van so Rosa can pick up the trade fair brochures on the 18th.
529 tokens in, 397 out (372 of them reasoning) · 6.3 s · $0.0058 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·
GPT-6 Astra
Passed: Graded 5.0 of 5 on average (lowest 5).
Paper shortages delayed brochures until the 18th; Rosa will collect them that day using the van booked for 8:00–12:00, ahead of the trade fair on the 20th.
483 tokens in, 45 out · 1.9 s · $0.0071 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·
More summarization prompts
The other summarization prompts, each with every model's reply, and the results across all five.
Questions
Which AI does best on “An email thread in one sentence”?
10 of 16 models passed this summarization prompt's check (rubric); Claude Fable 5.1, Claude Sonnet 5.5, Claude Sonnet 5, DeepSeek V4 Pro, Kimi K3, and GLM 5.3 Flash didn't. The cheapest reply that passed was GPT-6 Luna's, at $0.000067; the fastest, GLM 5.3's in 1.0 s. The dearest reply, Claude Fable 5.1's, cost 170 times as much ($0.0114).
What does a reply to “An email thread in one sentence” cost?
Through the models' APIs, what OpenRouter charged us ran from $0.000067 (GPT-6 Luna) to $0.0114 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.