Skip to content

Tested prompt · Summarization

A quarterly memo for the CEO: every AI model's reply, tested

We sent this hard summarization prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.

Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

All 16 models passed this summarization prompt's check (rubric). The cheapest reply that passed was DeepSeek V4.1 Flash's, at $0.000070; the fastest, GLM 5.3's in 1.1 s. The dearest reply, Claude Fable 5.1's, cost 265 times as much ($0.0184).

The prompt, as sent, and its check

Checked by rubric (graded), the same way for every model.

A quarterly memo for the CEO (hard)

Summarize this memo for the CEO in at most 4 sentences, keeping every number exact.

Q3 at Fernway: revenue was $2.4M, up 18% on Q2, driven by the new team plan, which now makes up 41% of new sign-ups. Churn fell from 3.1% to 2.6% a month. Gross margin slipped from 71% to 68% because of higher hosting costs, which the platform team expects to bring back down by the end of Q4 after moving storage providers. We ended the quarter with 14 months of runway. Hiring is frozen until Q2 next year, except for two support roles already approved. The biggest risk is our largest customer, 9% of revenue, whose contract is up for renewal in November.

Graded on:

  • Right priorities: Leads with what a CEO needs: growth, churn, margin, runway, the renewal risk.
  • Exact numbers: Every number exact and correctly attributed; nothing invented.
  • Crisp: Reads as a crisp executive summary.

Exactly what this prompt's replies are checked against, with every other prompt of our test runs.

Every model's result

All 16 models on this prompt, in catalog order.

Every model's reply to “A quarterly memo for the CEO”
ModelResultCostTimeReply
Claude Fable 5.1AnthropicPassed: Graded 4.7 of 5 on average (lowest 4).$0.01845.0 s197 tokens
Claude Opus 5.5AnthropicPassed: Graded 5.0 of 5 on average (lowest 5).$0.00753.5 s206 tokens
Claude Sonnet 5.5AnthropicPassed: Graded 4.7 of 5 on average (lowest 4).$0.00371.8 s195 tokens
Claude Sonnet 5AnthropicPassed: Graded 4.7 of 5 on average (lowest 4).$0.00333.0 s187 tokens
Claude Haiku 4.5AnthropicPassed: Graded 4.7 of 5 on average (lowest 4).$0.00121.5 s133 tokens
GPT-6 AstraOpenAIPassed: Graded 4.7 of 5 on average (lowest 4).$0.01182.9 s134 tokens
GPT-6 SolOpenAIPassed: Graded 4.7 of 5 on average (lowest 4).$0.00242.6 s141 tokens
GPT-6 LunaOpenAIPassed: Graded 4.0 of 5 on average (lowest 4).$0.000121.5 s141 tokens
Gemini 3.1 Pro (preview)GooglePassed: Graded 4.3 of 5 on average (lowest 4).$0.013911.2 s132 tokens
Gemini 3.8 FlashGooglePassed: Graded 4.7 of 5 on average (lowest 4).$0.000914.0 s136 tokens
DeepSeek V4.1 FlashDeepSeekPassed: Graded 4.7 of 5 on average (lowest 4).$0.0000703.2 s135 tokens
DeepSeek V4 ProDeepSeekPassed: Graded 4.3 of 5 on average (lowest 4).$0.000183.1 s125 tokens
Grok 4.7xAIPassed: Graded 4.0 of 5 on average (lowest 3).$0.00446.1 s139 tokens
Kimi K3MoonshotPassed: Graded 5.0 of 5 on average (lowest 5).$0.004915.7 s139 tokens
GLM 5.3Z.aiPassed: Graded 4.3 of 5 on average (lowest 4).$0.000991.1 s133 tokens
GLM 5.3 FlashZ.aiPassed: Graded 4.7 of 5 on average (lowest 4).$0.000158.3 s140 tokens

Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.

Every reply

Every reply passed: here they are from the cheapest up.

  1. DeepSeek V4.1 Flash

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Q3 revenue hit $2.4M, up 18% on Q2, driven by the new team plan that now accounts for 41% of new sign-ups, while monthly churn improved from 3.1% to 2.6%. Gross margin fell from 71% to 68% on higher hosting costs, which the platform team expects to restore by end of Q4 after switching storage providers. We closed the quarter with 14 months of runway, hiring is frozen until Q2 next year except for two already-approved support roles. The top risk is our largest customer at 9% of revenue, whose contract renews in November.

    545 tokens in, 146 out (11 of them reasoning) · 3.2 s · $0.0001 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Morph ·

  2. GPT-6 Luna

    Passed: Graded 4.0 of 5 on average (lowest 4).

    Q3 revenue was $2.4M, up 18% on Q2, with the new team plan driving growth and accounting for 41% of new sign-ups; monthly churn fell from 3.1% to 2.6%. Gross margin declined from 71% to 68% due to higher hosting costs, which the platform team expects to reduce by the end of Q4 after a storage-provider move. Fernway ended the quarter with 14 months of runway, and hiring is frozen until Q2 next year except for two approved support roles. The main risk is the largest customer, representing 9% of revenue, whose contract renews in November.

    507 tokens in, 141 out · 1.5 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  3. GLM 5.3 Flash

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Fernway's Q3 revenue hit $2.4M, up 18% on Q2, with the new team plan driving 41% of new sign-ups and monthly churn improving from 3.1% to 2.6%. Gross margin slipped from 71% to 68% due to higher hosting costs, but the platform team expects to reverse this by end of Q4 after switching storage providers. The company ended the quarter with 14 months of runway, and hiring is frozen until Q2 next year, except for two approved support roles. The biggest risk is renewal in November of our largest customer's contract, which accounts for 9% of revenue.

    522 tokens in, 140 out · 8.3 s · $0.0001 · 1 message on Pro · answered by z-ai/glm-5.3-flash via Near AI ·

  4. DeepSeek V4 Pro

    Passed: Graded 4.3 of 5 on average (lowest 4).

    Q3 revenue reached $2.4M, up 18% on Q2, driven by the new team plan, now 41% of new sign-ups. Monthly churn improved from 3.1% to 2.6%, but gross margin slipped from 71% to 68% on higher hosting costs, which the platform team expects to restore by end of Q4. We hold 14 months of runway, with hiring frozen until Q2 next year except two approved support roles. The biggest risk is our largest customer, 9% of revenue, whose contract renews in November.

    517 tokens in, 153 out (28 of them reasoning) · 3.1 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Baidu ·

  5. Gemini 3.8 Flash

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Q3 revenue reached $2.4M (up 18% on Q2), propelled by the new team plan accounting for 41% of new sign-ups, while monthly churn dropped from 3.1% to 2.6%. Gross margin declined from 71% to 68% due to hosting costs, though a storage migration is expected to resolve this by the end of Q4. The company holds 14 months of runway, with hiring frozen until Q2 next year aside from two approved support roles. The primary risk is the upcoming November renewal for Fernway's largest customer, which represents 9% of revenue.

    537 tokens in, 136 out · 4.0 s · $0.0009 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

  6. GLM 5.3

    Passed: Graded 4.3 of 5 on average (lowest 4).

    Q3 revenue was $2.4M, up 18% on Q2, driven by the new team plan (41% of new sign-ups), while monthly churn fell from 3.1% to 2.6%. Gross margin slipped from 71% to 68% due to higher hosting costs, but the platform team expects it back down by end of Q4 after moving storage providers. We ended the quarter with 14 months of runway, with hiring frozen until Q2 next year except for two approved support roles. The biggest risk is our largest customer, 9% of revenue, whose contract is up for renewal in November.

    545 tokens in, 134 out (1 of them reasoning) · 1.1 s · $0.0010 · 1 message on Pro · answered by z-ai/glm-5.3 via Wafer ·

  7. Claude Haiku 4.5

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Q3 revenue reached $2.4M, up 18% from Q2, with the new team plan driving 41% of new sign-ups and monthly churn improving from 3.1% to 2.6%. Gross margin declined from 71% to 68% due to higher hosting costs, expected to recover by end of Q4 after a storage provider migration. The company has 14 months of runway with hiring frozen until Q2 next year except for two approved support roles. Key risk: the largest customer, representing 9% of revenue, has a contract renewal due in November.

    543 tokens in, 133 out · 1.5 s · $0.0012 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

  8. GPT-6 Sol

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Fernway’s Q3 revenue was $2.4M, up 18% on Q2, driven by the new team plan, which accounts for 41% of new sign-ups. Monthly churn fell from 3.1% to 2.6%, while gross margin slipped from 71% to 68% due to higher hosting costs; the platform team expects to bring those costs back down by the end of Q4 after moving storage providers. Runway stood at 14 months, and hiring is frozen until Q2 next year except for two approved support roles. The biggest risk is the November renewal of the largest customer, which represents 9% of revenue.

    507 tokens in, 141 out · 2.6 s · $0.0024 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  9. Claude Sonnet 5

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Q3 revenue reached $2.4M, up 18% from Q2, driven by the team plan (now 41% of new sign-ups), while monthly churn improved from 3.1% to 2.6%. Gross margin dipped from 71% to 68% due to higher hosting costs, but the platform team expects to restore margins by end of Q4 after switching storage providers. Runway stands at 14 months, and hiring remains frozen until Q2 next year except for two already-approved support roles. Key risk: our largest customer (9% of revenue) has a contract renewal due in November.

    693 tokens in, 187 out · 3.0 s · $0.0033 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

  10. Claude Sonnet 5.5

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Fernway's Q3 revenue was $2.4M, up 18% on Q2, driven by the new team plan (now 41% of new sign-ups), while monthly churn fell from 3.1% to 2.6%. Gross margin slipped from 71% to 68% due to higher hosting costs, which the platform team expects to reverse by the end of Q4 after moving storage providers. Runway stands at 14 months, and hiring is frozen until Q2 next year except for two already-approved support roles. The biggest risk is the largest customer (9% of revenue), whose contract is up for renewal in November.

    697 tokens in, 195 out · 1.8 s · $0.0037 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·

  11. Grok 4.7

    Passed: Graded 4.0 of 5 on average (lowest 3).

    Q3 revenue at Fernway was $2.4M, up 18% from Q2, with the new team plan making up 41% of new sign-ups and monthly churn falling from 3.1% to 2.6%. Gross margin slipped from 71% to 68% on higher hosting costs, which the platform team expects to bring back down by the end of Q4 after moving storage providers. The quarter closed with 14 months of runway, and hiring is frozen until Q2 next year except for two already-approved support roles. The biggest risk is the largest customer, 9% of revenue, whose contract is up for renewal in November.

    1,741 tokens in, 615 out (476 of them reasoning) · 6.1 s · $0.0044 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  12. Kimi K3

    Passed: Graded 5.0 of 5 on average (lowest 5).

    Q3 revenue was $2.4M, up 18% on Q2, with the new team plan now 41% of new sign-ups. Churn improved from 3.1% to 2.6% a month, but gross margin slipped from 71% to 68% on higher hosting costs, which platform expects to fix by end of Q4 after moving storage providers. Fernway ended the quarter with 14 months of runway; hiring is frozen until Q2 next year except for two approved support roles. Key risk: the largest customer is 9% of revenue and its contract renews in November.

    593 tokens in, 506 out (367 of them reasoning) · 15.7 s · $0.0049 · 1 message on Pro · answered by moonshotai/kimi-k3 via Wafer ·

  13. Claude Opus 5.5

    Passed: Graded 5.0 of 5 on average (lowest 5).

    Fernway's Q3 revenue was $2.4M, up 18% on Q2, driven by the new team plan, which now accounts for 41% of new sign-ups, while monthly churn fell from 3.1% to 2.6%. Gross margin slipped from 71% to 68% due to higher hosting costs, which the platform team expects to bring down by the end of Q4 after switching storage providers. We ended the quarter with 14 months of runway, and hiring is frozen until Q2 next year except for two already-approved support roles. The biggest risk is our largest customer, 9% of revenue, whose contract is up for renewal in November.

    695 tokens in, 206 out · 3.5 s · $0.0075 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  14. GPT-6 Astra

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Q3 revenue was $2.4M, up 18% on Q2, driven by the new team plan, which accounts for 41% of new sign-ups; monthly churn fell from 3.1% to 2.6%. Gross margin declined from 71% to 68% due to higher hosting costs, which the platform team expects to reduce by the end of Q4 after switching storage providers. Quarter-end runway was 14 months, and hiring is frozen until Q2 next year except for two approved support roles. The biggest risk is the November renewal of our largest customer, which represents 9% of revenue.

    507 tokens in, 134 out · 2.9 s · $0.0118 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

  15. Gemini 3.1 Pro

    Passed: Graded 4.3 of 5 on average (lowest 4).

    In Q3, revenue reached $2.4M (up 18% from Q2), driven by the new team plan that now constitutes 41% of new sign-ups. Monthly churn improved from 3.1% to 2.6%, though higher hosting costs temporarily reduced gross margin from 71% to 68% until a planned Q4 provider switch. We currently maintain 14 months of runway and have frozen hiring until Q2 of next year, excluding 2 already-approved support roles. Our primary risk is the upcoming November renewal for our largest customer, who accounts for 9% of revenue.

    537 tokens in, 1,071 out (939 of them reasoning) · 11.2 s · $0.0139 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  16. Claude Fable 5.1

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Fernway's Q3 revenue was $2.4M, up 18% on Q2, driven by the new team plan (41% of new sign-ups), while monthly churn fell from 3.1% to 2.6%. Gross margin slipped from 71% to 68% due to higher hosting costs, which the platform team expects to recover by end of Q4 after switching storage providers. We ended the quarter with 14 months of runway, and hiring is frozen until Q2 next year except for two already-approved support roles. The biggest risk is our largest customer (9% of revenue), whose contract comes up for renewal in November.

    695 tokens in, 197 out · 5.0 s · $0.0184 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

More summarization prompts

The other summarization prompts, each with every model's reply, and the results across all five.

Questions

Which AI does best on “A quarterly memo for the CEO”?

All 16 models passed this summarization prompt's check (rubric). The cheapest reply that passed was DeepSeek V4.1 Flash's, at $0.000070; the fastest, GLM 5.3's in 1.1 s. The dearest reply, Claude Fable 5.1's, cost 265 times as much ($0.0184).

What does a reply to “A quarterly memo for the CEO” cost?

Through the models' APIs, what OpenRouter charged us ran from $0.000070 (DeepSeek V4.1 Flash) to $0.0184 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.