Skip to content

Tested prompt · Data analysis

A median, filtered two ways: every AI model's reply, tested

We sent this hard data analysis prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.

Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

All 16 models passed this data analysis prompt's check (final answer). The cheapest reply that passed was GPT-6 Luna's, at $0.00012; the fastest, GLM 5.3's in 1.0 s. The dearest reply, Claude Fable 5.1's, cost 131 times as much ($0.0151).

The prompt, as sent, and its check

Checked by final answer, the same way for every model.

A median, filtered two ways (hard)

Here are support tickets as CSV. What's the median hours_to_resolve for high-priority tickets handled by the Billing team?

ticket_id,team,priority,hours_to_resolve
T1,Billing,high,5.5
T2,Tech,high,12
T3,Billing,low,30
T4,Billing,high,2
T5,Tech,low,48
T6,Billing,high,9
T7,Billing,medium,16
T8,Tech,high,3.5
T9,Billing,high,7
T10,Billing,high,26
T11,Tech,medium,20
T12,Billing,high,4

Show your working briefly, then end with a line that says "Final answer: " followed by the answer alone.

The final answer must be 6.25.

Exactly what this prompt's replies are checked against, with every other prompt of our test runs.

Every model's result

All 16 models on this prompt, in catalog order.

Every model's reply to “A median, filtered two ways”
ModelResultCostTimeReply
Claude Fable 5.1AnthropicPassed: Final answer 6.25: right.$0.01516.7 s132 tokens
Claude Opus 5.5AnthropicPassed: Final answer 6.25: right.$0.00794.6 s142 tokens
Claude Sonnet 5.5AnthropicPassed: Final answer 6.25 hours: right.$0.00301.5 s127 tokens
Claude Sonnet 5AnthropicPassed: Final answer 6.25 hours: right.$0.00364.9 s123 tokens
Claude Haiku 4.5AnthropicPassed: Final answer 6.25: right.$0.00162.0 s216 tokens
GPT-6 AstraOpenAIPassed: Final answer 6.25 hours: right.$0.00902.0 s77 tokens
GPT-6 SolOpenAIPassed: Final answer 6.25 hours: right.$0.00242.5 s94 tokens
GPT-6 LunaOpenAIPassed: Final answer 6.25 hours: right.$0.000122.5 s75 tokens
Gemini 3.1 Pro (preview)GooglePassed: Final answer 6.25: right.$0.00827.1 s185 tokens
Gemini 3.8 FlashGooglePassed: Final answer 6.25: right.$0.00103.5 s169 tokens
DeepSeek V4.1 FlashDeepSeekPassed: Final answer 6.25: right.$0.000251.6 s107 tokens
DeepSeek V4 ProDeepSeekPassed: Final answer 6.25: right.$0.000851.3 s77 tokens
Grok 4.7xAIPassed: Final answer 6.25: right.$0.00364.2 s88 tokens
Kimi K3MoonshotPassed: Final answer 6.25: right.$0.00423.5 s83 tokens
GLM 5.3Z.aiPassed: Final answer 6.25: right.$0.00111.0 s92 tokens
GLM 5.3 FlashZ.aiPassed: Final answer 6.25: right.$0.000123.7 s89 tokens

Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.

Every reply

Every reply passed: here they are from the cheapest up.

  1. GPT-6 Luna

    Passed: Final answer 6.25 hours: right.

    Billing high-priority resolution times, sorted: 2, 4, 5.5, 7, 9, 26. With six values, the median is the average of the middle two: (5.5 + 7) / 2 = 6.25 hours.
    
    Final answer: 6.25 hours

    512 tokens in, 128 out (53 of them reasoning) · 2.5 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  2. GLM 5.3 Flash

    Passed: Final answer 6.25: right.

    Billing + high tickets: T1=5.5, T4=2, T6=9, T9=7, T10=26, T12=4
    
    Sorted: 2, 4, 5.5, 7, 9, 26 — six values, so median = (5.5 + 7)/2 = 6.25
    
    Final answer: 6.25

    526 tokens in, 142 out (53 of them reasoning) · 3.7 s · $0.0001 · 1 message on Pro · answered by z-ai/glm-5.3-flash via AtlasCloud ·

  3. DeepSeek V4.1 Flash

    Passed: Final answer 6.25: right.

    **Filter:** team = Billing AND priority = high
    
    | ticket | hours |
    |---|---|
    | T4 | 2 |
    | T12 | 4 |
    | T1 | 5.5 |
    | T9 | 7 |
    | T6 | 9 |
    | T10 | 26 |
    
    That's 6 values; median = mean of the 3rd and 4th = (5.5 + 7) / 2 = **6.25**
    
    Final answer: 6.25

    554 tokens in, 209 out (102 of them reasoning) · 1.6 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Novita ·

  4. DeepSeek V4 Pro

    Passed: Final answer 6.25: right.

    Filtering Billing + high priority: 5.5, 2, 9, 7, 26, 4  
    Sorted: 2, 4, 5.5, 7, 9, 26  
    Median = (5.5 + 7) / 2 = 6.25
    
    Final answer: 6.25

    526 tokens in, 205 out (128 of them reasoning) · 1.3 s · $0.0008 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Wafer ·

  5. Gemini 3.8 Flash

    Passed: Final answer 6.25: right.

    1. Filter for tickets where `team` is "Billing" and `priority` is "high":
       - T1: 5.5
       - T4: 2
       - T6: 9
       - T9: 7
       - T10: 26
       - T12: 4
    
    2. Sort the values in ascending order:
       - 2, 4, 5.5, 7, 9, 26
    
    3. Calculate the median of the 6 values (average of the 3rd and 4th values):
       - (5.5 + 7) / 2 = 12.5 / 2 = 6.25
    
    Final answer: 6.25

    551 tokens in, 169 out · 3.5 s · $0.0010 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

  6. GLM 5.3

    Passed: Final answer 6.25: right.

    Billing high-priority tickets: 5.5, 2, 9, 7, 26, 4 (six values).
    
    Sorted: 2, 4, 5.5, 7, 9, 26. With an even count, the median is the average of the two middle values: (5.5 + 7) / 2 = 6.25.
    
    Final answer: 6.25

    549 tokens in, 149 out (57 of them reasoning) · 1.0 s · $0.0011 · 1 message on Pro · answered by z-ai/glm-5.3 via Wafer ·

  7. Claude Haiku 4.5

    Passed: Final answer 6.25: right.

    I'll filter for high-priority Billing team tickets and find the median.
    
    **Filtering the data:**
    
    High-priority tickets handled by Billing team:
    - T1: 5.5 hours
    - T4: 2 hours
    - T6: 9 hours
    - T9: 7 hours
    - T10: 26 hours
    - T12: 4 hours
    
    **Sorting by hours_to_resolve:**
    2, 4, 5.5, 7, 9, 26
    
    **Calculating median:**
    With 6 values (even number), the median is the average of the 3rd and 4th values:
    - 3rd value: 5.5
    - 4th value: 7
    - Median: (5.5 + 7) / 2 = 12.5 / 2 = 6.25
    
    Final answer: 6.25

    550 tokens in, 216 out · 2.0 s · $0.0016 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

  8. GPT-6 Sol

    Passed: Final answer 6.25 hours: right.

    The high-priority Billing tickets took 5.5, 2, 9, 7, 26, and 4 hours. Sorted: 2, 4, 5.5, 7, 9, 26. The median is the average of the middle two values: (5.5 + 7) ÷ 2 = 6.25.
    
    Final answer: 6.25 hours

    512 tokens in, 137 out (43 of them reasoning) · 2.5 s · $0.0024 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  9. Claude Sonnet 5.5

    Passed: Final answer 6.25 hours: right.

    Billing + high tickets: T1 (5.5), T4 (2), T6 (9), T9 (7), T10 (26), T12 (4).
    
    Sorted: 2, 4, 5.5, 7, 9, 26
    
    Six values, so the median is the average of the 3rd and 4th: (5.5 + 7) / 2 = 6.25
    
    Final answer: 6.25 hours

    693 tokens in, 127 out · 1.5 s · $0.0030 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·

  10. Grok 4.7

    Passed: Final answer 6.25: right.

    High-priority Billing tickets: 5.5, 2, 9, 7, 26, 4.
    
    Sorted: 2, 4, 5.5, 7, 9, 26.
    
    Median of 6 values = average of 3rd and 4th = (5.5 + 7) / 2 = 6.25.
    
    Final answer: 6.25

    1,749 tokens in, 448 out (360 of them reasoning) · 4.2 s · $0.0036 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  11. Claude Sonnet 5

    Passed: Final answer 6.25 hours: right.

    Billing + high-priority tickets: T1=5.5, T4=2, T6=9, T9=7, T10=26, T12=4
    
    Sorted: 2, 4, 5.5, 7, 9, 26 (6 values)
    
    Median = average of 3rd and 4th values = (5.5 + 7) / 2 = 6.25
    
    Final answer: 6.25 hours

    689 tokens in, 223 out (100 of them reasoning) · 4.9 s · $0.0036 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

  12. Kimi K3

    Passed: Final answer 6.25: right.

    Billing + high tickets: 5.5, 2, 9, 7, 26, 4  
    Sorted: 2, 4, 5.5, 7, 9, 26  
    Median = (5.5 + 7) / 2 = 6.25  
    
    Final answer: 6.25

    597 tokens in, 265 out (182 of them reasoning) · 3.5 s · $0.0042 · 1 message on Pro · answered by moonshotai/kimi-k3 via Phala ·

  13. Claude Opus 5.5

    Passed: Final answer 6.25: right.

    **Billing + high-priority tickets:**
    - T1: 5.5
    - T4: 2
    - T6: 9
    - T9: 7
    - T10: 26
    - T12: 4
    
    **Sorted:** 2, 4, 5.5, 7, 9, 26
    
    There are 6 values, so the median is the average of the 3rd and 4th values: (5.5 + 7) / 2 = **6.25**
    
    Final answer: 6.25

    691 tokens in, 225 out (83 of them reasoning) · 4.6 s · $0.0079 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  14. Gemini 3.1 Pro

    Passed: Final answer 6.25: right.

    To find the median `hours_to_resolve` for high-priority tickets handled by the Billing team, we first filter the data for those specific tickets:
    
    *   T1: 5.5
    *   T4: 2
    *   T6: 9
    *   T9: 7
    *   T10: 26
    *   T12: 4
    
    Next, we sort these values in ascending order:
    2, 4, 5.5, 7, 9, 26
    
    Since there are 6 values (an even number), the median is the average of the two middle numbers (the 3rd and 4th values):
    (5.5 + 7) / 2 = 12.5 / 2 = 6.25
    
    Final answer: 6.25

    551 tokens in, 590 out (405 of them reasoning) · 7.1 s · $0.0082 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  15. GPT-6 Astra

    Passed: Final answer 6.25 hours: right.

    High-priority Billing resolution times, sorted: 2, 4, 5.5, 7, 9, 26 hours.
    
    With 6 tickets, the median is the average of the two middle values: (5.5 + 7) ÷ 2 = 6.25 hours.
    
    Final answer: 6.25 hours

    512 tokens in, 77 out · 2.0 s · $0.0090 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

  16. Claude Fable 5.1

    Passed: Final answer 6.25: right.

    Billing + high tickets: T1 (5.5), T4 (2), T6 (9), T9 (7), T10 (26), T12 (4) — 6 tickets.
    
    Sorted hours: 2, 4, 5.5, 7, 9, 26
    
    Median of 6 values = average of 3rd and 4th: (5.5 + 7) / 2 = 6.25
    
    Final answer: 6.25

    691 tokens in, 132 out · 6.7 s · $0.0151 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

More data analysis prompts

The other data analysis prompts, each with every model's reply.

Questions

Which AI does best on “A median, filtered two ways”?

All 16 models passed this data analysis prompt's check (final answer). The cheapest reply that passed was GPT-6 Luna's, at $0.00012; the fastest, GLM 5.3's in 1.0 s. The dearest reply, Claude Fable 5.1's, cost 131 times as much ($0.0151).

What does a reply to “A median, filtered two ways” cost?

Through the models' APIs, what OpenRouter charged us ran from $0.00012 (GPT-6 Luna) to $0.0151 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.