Skip to content

Tested prompt · Writing

Rewrite corporate jargon in plain words: every AI model's reply, tested

We sent this everyday writing prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.

Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

11 of 16 models passed this writing prompt's check (rubric); GPT-6 Sol, Gemini 3.8 Flash, DeepSeek V4.1 Flash, Grok 4.7, and GLM 5.3 Flash didn't. The cheapest reply that passed was GPT-6 Luna's, at $0.00010; the fastest, DeepSeek V4 Pro's in 1.2 s. The dearest reply, GPT-6 Astra's, cost 140 times as much ($0.0145).

The prompt, as sent, and its check

Checked by rubric (graded), the same way for every model.

Rewrite corporate jargon in plain words (everyday)

Rewrite this for a general audience in at most 90 words, keeping every fact:

"Leveraging our cross-functional synergies, the Q3 initiative operationalized a customer-centric paradigm shift, yielding a 12% uplift in retention KPIs and a 30-basis-point reduction in churn velocity across the enterprise segment, while our omnichannel touchpoints were right-sized to optimize bandwidth."

Graded on:

  • Plain words: Plain words a general reader understands; no jargon left.
  • Keeps the facts: Keeps the facts: a Q3 project, 12% better retention, churn down 0.3 percentage points among large business customers, fewer or better-sized support channels.
  • Clear: Short and clear.

Exactly what this prompt's replies are checked against, with every other prompt of our test runs.

Every model's result

All 16 models on this prompt, in catalog order.

Every model's reply to “Rewrite corporate jargon in plain words”
ModelResultCostTimeReply
Claude Fable 5.1AnthropicPassed: Graded 4.3 of 5 on average (lowest 4).$0.01306.0 s114 tokens
Claude Opus 5.5AnthropicPassed: Graded 4.3 of 5 on average (lowest 4).$0.00785.1 s112 tokens
Claude Sonnet 5.5AnthropicPassed: Graded 4.7 of 5 on average (lowest 4).$0.00251.6 s102 tokens
Claude Sonnet 5AnthropicPassed: Graded 4.7 of 5 on average (lowest 4).$0.00222.7 s105 tokens
Claude Haiku 4.5AnthropicPassed: Graded 4.3 of 5 on average (lowest 4).$0.000831.6 s73 tokens
GPT-6 AstraOpenAIPassed: Graded 4.0 of 5 on average (lowest 4).$0.01456.0 s72 tokens
GPT-6 SolOpenAIFailed: Graded 3.7 of 5 on average (lowest 3).$0.00335.0 s77 tokens
GPT-6 LunaOpenAIPassed: Graded 4.7 of 5 on average (lowest 4).$0.000102.4 s65 tokens
Gemini 3.1 Pro (preview)GooglePassed: Graded 4.0 of 5 on average (lowest 3).$0.01018.8 s76 tokens
Gemini 3.8 FlashGoogleFailed: Graded 3.7 of 5 on average (lowest 3).$0.000643.1 s81 tokens
DeepSeek V4.1 FlashDeepSeekFailed: Graded 3.7 of 5 on average (lowest 3).$0.000291.2 s68 tokens
DeepSeek V4 ProDeepSeekPassed: Graded 4.0 of 5 on average (lowest 3).$0.000591.2 s54 tokens
Grok 4.7xAIFailed: Graded 3.7 of 5 on average (lowest 3).$0.008819.3 s60 tokens
Kimi K3MoonshotPassed: Graded 4.3 of 5 on average (lowest 4).$0.00302.0 s82 tokens
GLM 5.3Z.aiPassed: Graded 4.0 of 5 on average (lowest 3).$0.000201.8 s57 tokens
GLM 5.3 FlashZ.aiFailed: Graded 3.3 of 5 on average (lowest 3).$0.0000974.9 s73 tokens

Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.

Every reply

The 5 replies that failed first, then the ones that passed, from the cheapest up.

  1. GLM 5.3 Flash

    Failed: Graded 3.3 of 5 on average (lowest 3).

    In Q3, teams from across the company worked together on an initiative that put customers first. The result: customer retention improved by 12%, and the rate at which enterprise customers left fell by 30 basis points. We also adjusted our customer contact channels—like phone, email, and chat—so each team could focus on what it does best.

    435 tokens in, 124 out (51 of them reasoning) · 4.9 s · $0.0001 · 1 message on Pro · answered by z-ai/glm-5.3-flash via AtlasCloud ·

  2. DeepSeek V4.1 Flash

    Failed: Graded 3.7 of 5 on average (lowest 3).

    In the third quarter, teams from across the company worked together on a project that put customers first. For our large business customers, retention rose 12%, and churn fell by 30 basis points. We also adjusted our channels — online, in-store, and other ways customers reach us — to use our resources more efficiently.

    459 tokens in, 129 out (61 of them reasoning) · 1.2 s · $0.0003 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Parasail ·

  3. Gemini 3.8 Flash

    Failed: Graded 3.7 of 5 on average (lowest 3).

    By having different teams work together in the third quarter, we refocused our business on customer needs. This effort paid off: client retention rates increased by 12%, and the rate at which enterprise clients canceled dropped by 0.30% (30 basis points). At the same time, we streamlined and balanced our communication channels to make better use of our team’s time and resources.

    447 tokens in, 81 out · 3.1 s · $0.0006 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

  4. GPT-6 Sol

    Failed: Graded 3.7 of 5 on average (lowest 3).

    In Q3, teams across the company worked together to put a more customer-focused approach into practice. In the enterprise customer segment, retention measures improved by 12%, while the rate at which customers left fell by 30 basis points (0.3 percentage points). We also adjusted how we connect with customers across channels to make better use of our capacity.

    417 tokens in, 246 out (169 of them reasoning) · 5.0 s · $0.0033 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  5. Grok 4.7

    Failed: Graded 3.7 of 5 on average (lowest 3).

    In the third quarter, teams from different departments worked together on an initiative that put a stronger customer focus into practice. Retention measures rose 12%, and the speed at which enterprise customers left fell by 30 basis points. We also resized customer contact points across channels so they used available capacity more efficiently.

    1,657 tokens in, 1,631 out (1,571 of them reasoning) · 19.3 s · $0.0088 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  6. GPT-6 Luna

    Passed: Graded 4.7 of 5 on average (lowest 4).

    In Q3, teams across the company worked together to focus more on customers. Retention measures improved by 12%, and the rate of customer loss fell by 0.3 percentage points in the enterprise segment. The company also adjusted its customer contact channels to make better use of available capacity.

    417 tokens in, 124 out (59 of them reasoning) · 2.4 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  7. GLM 5.3

    Passed: Graded 4.0 of 5 on average (lowest 3).

    In Q3, our teams worked together to make the business more customer-focused. As a result, customer retention rose 12%, and churn in our enterprise segment fell by 0.3 percentage points. We also adjusted our customer service channels to use our resources more efficiently.

    434 tokens in, 98 out (41 of them reasoning) · 1.8 s · $0.0002 · 1 message on Pro · answered by z-ai/glm-5.3 via Baidu ·

  8. DeepSeek V4 Pro

    Passed: Graded 4.0 of 5 on average (lowest 3).

    In the third quarter, our teams worked together to put customers first. As a result, retention among enterprise customers rose by 12%, and churn fell by 0.30 percentage points. We also adjusted our support and communication channels to use resources more efficiently.

    431 tokens in, 139 out (85 of them reasoning) · 1.2 s · $0.0006 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Wafer ·

  9. Claude Haiku 4.5

    Passed: Graded 4.3 of 5 on average (lowest 4).

    Here's a clearer version:
    
    Our Q3 initiative brought different departments together to focus on customer needs. The results were strong: we kept 12% more customers and reduced customer loss by 0.3 percentage points among large businesses. We also streamlined our customer contact channels to work more efficiently.
    
    (77 words)

    462 tokens in, 73 out · 1.6 s · $0.0008 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

  10. Claude Sonnet 5

    Passed: Graded 4.7 of 5 on average (lowest 4).

    In Q3, teams worked together to focus more on customers' needs. This effort increased customer retention by 12% and reduced the rate of customer loss (churn) by 0.3 percentage points among large business clients. The company also adjusted the number of customer contact channels (like phone, email, and chat) to better manage workload and resources.

    597 tokens in, 105 out · 2.7 s · $0.0022 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

  11. Claude Sonnet 5.5

    Passed: Graded 4.7 of 5 on average (lowest 4).

    In Q3, teams from different departments worked together to put customers first. The result was a 12% increase in customer retention and a 0.30 percentage-point drop in the rate at which enterprise customers leave. We also reduced the number of ways we contact customers (such as email, phone, and chat) to free up staff time and resources.

    601 tokens in, 102 out · 1.6 s · $0.0025 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·

  12. Kimi K3

    Passed: Graded 4.3 of 5 on average (lowest 4).

    In Q3, teams across the company worked together to focus more on customers. As a result, customer retention improved by 12%, and the rate at which enterprise customers left dropped by 30 basis points (0.3%). We also streamlined our communication channels—such as phone, email, and online—to use our time and resources more efficiently.

    504 tokens in, 120 out (38 of them reasoning) · 2.0 s · $0.0030 · 1 message on Pro · answered by moonshotai/kimi-k3 via Together ·

  13. Claude Opus 5.5

    Passed: Graded 4.3 of 5 on average (lowest 4).

    In the third quarter, teams from different departments worked together on a project to put customers at the center of how the company operates. As a result, customer retention measures rose 12%, and the rate at which large business customers left dropped by 0.3 percentage points. The company also streamlined its customer contact channels, including phone, email, and online, to make better use of staff time and resources.

    599 tokens in, 245 out (133 of them reasoning) · 5.1 s · $0.0078 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  14. Gemini 3.1 Pro

    Passed: Graded 4.0 of 5 on average (lowest 3).

    In Q3, our teams worked together to adopt a more customer-focused approach. For our business clients, this resulted in a 12% increase in customer retention and a 0.30% (30 basis points) decrease in the rate of customers leaving. Additionally, we adjusted our various customer communication channels to make more efficient use of our time and resources.

    447 tokens in, 763 out (687 of them reasoning) · 8.8 s · $0.0101 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  15. Claude Fable 5.1

    Passed: Graded 4.3 of 5 on average (lowest 4).

    In Q3, teams from different departments worked together on a project to put customers first. Among enterprise (large business) customers, this raised the retention rate by 12% and slowed the rate at which customers left by 0.3 percentage points. The company also cut back on some of the ways it communicates with customers—across channels like phone, email, and in-app—to free up staff time.

    599 tokens in, 114 out · 6.0 s · $0.0130 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

  16. GPT-6 Astra

    Passed: Graded 4.0 of 5 on average (lowest 4).

    In Q3, teams across departments worked together to shift toward a more customer-focused approach. Among enterprise customers, the initiative increased retention by 12% and reduced the rate at which customers leave by 0.3 percentage points (30 basis points). We also adjusted customer interactions across all channels to make better use of available capacity.

    417 tokens in, 206 out (134 of them reasoning) · 6.0 s · $0.0145 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

More writing prompts

The other writing prompts, each with every model's reply, and the results across all five.

Questions

Which AI does best on “Rewrite corporate jargon in plain words”?

11 of 16 models passed this writing prompt's check (rubric); GPT-6 Sol, Gemini 3.8 Flash, DeepSeek V4.1 Flash, Grok 4.7, and GLM 5.3 Flash didn't. The cheapest reply that passed was GPT-6 Luna's, at $0.00010; the fastest, DeepSeek V4 Pro's in 1.2 s. The dearest reply, GPT-6 Astra's, cost 140 times as much ($0.0145).

What does a reply to “Rewrite corporate jargon in plain words” cost?

Through the models' APIs, what OpenRouter charged us ran from $0.000097 (GLM 5.3 Flash) to $0.0145 (GPT-6 Astra) for this prompt. The cheapest that passed was GPT-6 Luna's, $0.00010. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.