Skip to content

Tested prompt · SQL

Count orders by status: every AI model's reply, tested

We sent this everyday SQL prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.

Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

All 16 models passed this SQL prompt's check (query result). The cheapest reply that passed was GPT-6 Luna's, at $0.000058; the fastest, Claude Haiku 4.5's in 0.8 s. The dearest reply, Claude Fable 5.1's, cost 210 times as much ($0.0122).

The prompt, as sent, and its check

Checked by query result, the same way for every model.

Count orders by status (everyday)

[5 lines every SQL prompt of ours shares, word for word: the whole prompt, on the methods page]
How many orders have the status 'shipped'?
Write one SQLite query and reply with it in a ```sql code block.

The query must return the same rows as this one, on the fixture database below:

SELECT COUNT(*) FROM orders WHERE status = 'shipped'

Exactly what this prompt's replies are checked against, with every other prompt of our test runs.

Every model's result

All 16 models on this prompt, in catalog order.

Every model's reply to “Count orders by status”
ModelResultCostTimeReply
Claude Fable 5.1AnthropicPassed: Returned the right 1 row.$0.01223.5 s46 tokens
Claude Opus 5.5AnthropicPassed: Returned the right 1 row.$0.00492.5 s46 tokens
Claude Sonnet 5.5AnthropicPassed: Returned the right 1 row.$0.00241.0 s46 tokens
Claude Sonnet 5AnthropicPassed: Returned the right 1 row.$0.00202.0 s36 tokens
Claude Haiku 4.5AnthropicPassed: Returned the right 1 row.$0.000630.8 s21 tokens
GPT-6 AstraOpenAIPassed: Returned the right 1 row.$0.00581.4 s20 tokens
GPT-6 SolOpenAIPassed: Returned the right 1 row.$0.00121.1 s25 tokens
GPT-6 LunaOpenAIPassed: Returned the right 1 row.$0.0000580.9 s20 tokens
Gemini 3.1 Pro (preview)GooglePassed: Returned the right 1 row.$0.00395.0 s17 tokens
Gemini 3.8 FlashGooglePassed: Returned the right 1 row.$0.000473.7 s21 tokens
DeepSeek V4.1 FlashDeepSeekPassed: Returned the right 1 row.$0.000121.0 s23 tokens
DeepSeek V4 ProDeepSeekPassed: Returned the right 1 row.$0.000192.7 s23 tokens
Grok 4.7xAIPassed: Returned the right 1 row.$0.00171.4 s17 tokens
Kimi K3MoonshotPassed: Returned the right 1 row.$0.000861.5 s31 tokens
GLM 5.3Z.aiPassed: Returned the right 1 row.$0.000131.2 s18 tokens
GLM 5.3 FlashZ.aiPassed: Returned the right 1 row.$0.000132.1 s17 tokens

Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.

Every reply

Every reply passed: here they are from the cheapest up.

  1. GPT-6 Luna

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) FROM orders WHERE status = 'shipped';
    ```

    481 tokens in, 20 out · 0.9 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  2. DeepSeek V4.1 Flash

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) AS shipped_orders
    FROM orders
    WHERE status = 'shipped';
    ```

    522 tokens in, 26 out (3 of them reasoning) · 1.0 s · $0.0001 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Novita ·

  3. GLM 5.3 Flash

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) FROM orders WHERE status = 'shipped';
    ```

    496 tokens in, 37 out (20 of them reasoning) · 2.1 s · $0.0001 · 1 message on Pro · answered by z-ai/glm-5.3-flash via Wafer ·

  4. GLM 5.3

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) FROM orders WHERE status = 'shipped';
    ```

    495 tokens in, 18 out · 1.2 s · $0.0001 · 1 message on Pro · answered by z-ai/glm-5.3 via Baidu ·

  5. DeepSeek V4 Pro

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) AS shipped_orders
    FROM orders
    WHERE status = 'shipped';
    ```

    494 tokens in, 75 out (52 of them reasoning) · 2.7 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via StreamLake ·

  6. Gemini 3.8 Flash

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) 
    FROM orders 
    WHERE status = 'shipped';
    ```

    520 tokens in, 21 out · 3.7 s · $0.0005 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

  7. Claude Haiku 4.5

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) FROM orders WHERE status = 'shipped';
    ```

    525 tokens in, 21 out · 0.8 s · $0.0006 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

  8. Kimi K3

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) AS shipped_orders
    FROM orders
    WHERE status = 'shipped';
    ```

    568 tokens in, 37 out (6 of them reasoning) · 1.5 s · $0.0009 · 1 message on Pro · answered by moonshotai/kimi-k3 via Wafer ·

  9. GPT-6 Sol

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) AS shipped_orders
    FROM orders
    WHERE status = 'shipped';
    ```

    481 tokens in, 25 out · 1.1 s · $0.0012 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  10. Grok 4.7

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) FROM orders WHERE status = 'shipped';
    ```

    1,727 tokens in, 61 out (44 of them reasoning) · 1.4 s · $0.0017 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  11. Claude Sonnet 5

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) FROM orders WHERE status = 'shipped';
    ```

    800 tokens in, 36 out · 2.0 s · $0.0020 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

  12. Claude Sonnet 5.5

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) AS shipped_orders
    FROM orders
    WHERE status = 'shipped';
    ```

    804 tokens in, 46 out · 1.0 s · $0.0024 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·

  13. Gemini 3.1 Pro

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) FROM orders WHERE status = 'shipped';
    ```

    520 tokens in, 237 out (220 of them reasoning) · 5.0 s · $0.0039 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  14. Claude Opus 5.5

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) AS shipped_orders
    FROM orders
    WHERE status = 'shipped';
    ```

    802 tokens in, 46 out · 2.5 s · $0.0049 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  15. GPT-6 Astra

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) FROM orders WHERE status = 'shipped';
    ```

    481 tokens in, 20 out · 1.4 s · $0.0058 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

  16. Claude Fable 5.1

    Passed: Returned the right 1 row.

    ```sql
    SELECT COUNT(*) AS shipped_orders
    FROM orders
    WHERE status = 'shipped';
    ```

    802 tokens in, 46 out · 3.5 s · $0.0122 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

More SQL prompts

The other SQL prompts, each with every model's reply, and the results across all five.

Questions

Which AI does best on “Count orders by status”?

All 16 models passed this SQL prompt's check (query result). The cheapest reply that passed was GPT-6 Luna's, at $0.000058; the fastest, Claude Haiku 4.5's in 0.8 s. The dearest reply, Claude Fable 5.1's, cost 210 times as much ($0.0122).

What does a reply to “Count orders by status” cost?

Through the models' APIs, what OpenRouter charged us ran from $0.000058 (GPT-6 Luna) to $0.0122 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.