Skip to content

Tested prompt · SQL

Revenue by category: every AI model's reply, tested

We sent this everyday SQL prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.

Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

All 16 models passed this SQL prompt's check (query result). The cheapest reply that passed was GPT-6 Luna's, at $0.000090, and the fastest too, in 1.2 s. The dearest reply, Claude Fable 5.1's, cost 210 times as much ($0.0190).

The prompt, as sent, and its check

Checked by query result, the same way for every model.

Revenue by category (everyday)

[5 lines every SQL prompt of ours shares, word for word: the whole prompt, on the methods page]
For completed orders only, show each product category and its revenue in dollars (quantity times price), highest revenue first.
Write one SQLite query and reply with it in a ```sql code block.

The query must return the same rows as this one, on the fixture database below, in the same order:

SELECT p.category, SUM(oi.quantity * p.price_cents) / 100.0 AS revenue
FROM orders o
JOIN order_items oi ON oi.order_id = o.id
JOIN products p ON p.id = oi.product_id
WHERE o.status = 'completed'
GROUP BY p.category
ORDER BY revenue DESC

Exactly what this prompt's replies are checked against, with every other prompt of our test runs.

Every model's result

All 16 models on this prompt, in catalog order.

Every model's reply to “Revenue by category”
ModelResultCostTimeReply
Claude Fable 5.1AnthropicPassed: Returned the right 3 rows.$0.01904.5 s177 tokens
Claude Opus 5.5AnthropicPassed: Returned the right 3 rows.$0.00924.3 s209 tokens
Claude Sonnet 5.5AnthropicPassed: Returned the right 3 rows.$0.00391.6 s187 tokens
Claude Sonnet 5AnthropicPassed: Returned the right 3 rows.$0.00313.5 s141 tokens
Claude Haiku 4.5AnthropicPassed: Returned the right 3 rows.$0.00111.2 s103 tokens
GPT-6 AstraOpenAIPassed: Returned the right 3 rows.$0.00912.2 s84 tokens
GPT-6 SolOpenAIPassed: Returned the right 3 rows.$0.00181.3 s84 tokens
GPT-6 LunaOpenAIPassed: Returned the right 3 rows.$0.0000901.2 s82 tokens
Gemini 3.1 Pro (preview)GooglePassed: Returned the right 3 rows.$0.00747.4 s96 tokens
Gemini 3.8 FlashGooglePassed: Returned the right 3 rows.$0.000763.0 s95 tokens
DeepSeek V4.1 FlashDeepSeekPassed: Returned the right 3 rows.$0.0000902.9 s98 tokens
DeepSeek V4 ProDeepSeekPassed: Returned the right 3 rows.$0.000203.8 s85 tokens
Grok 4.7xAIPassed: Returned the right 3 rows.$0.00314.9 s81 tokens
Kimi K3MoonshotPassed: Returned the right 3 rows.$0.00191.6 s84 tokens
GLM 5.3Z.aiPassed: Returned the right 3 rows.$0.000211.2 s79 tokens
GLM 5.3 FlashZ.aiPassed: Returned the right 3 rows.$0.000112.6 s78 tokens

Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.

Every reply

Every reply passed: here they are from the cheapest up.

  1. GPT-6 Luna

    Passed: Returned the right 3 rows.

    ```sql
    SELECT p.category,
           SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM orders AS o
    JOIN order_items AS oi ON oi.order_id = o.id
    JOIN products AS p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```

    493 tokens in, 82 out · 1.2 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  2. DeepSeek V4.1 Flash

    Passed: Returned the right 3 rows.

    ```sql
    SELECT p.category,
           ROUND(SUM(oi.quantity * p.price_cents) / 100.0, 2) AS revenue_dollars
    FROM order_items oi
    JOIN orders o ON o.id = oi.order_id
    JOIN products p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY SUM(oi.quantity * p.price_cents) DESC;
    ```

    534 tokens in, 106 out (8 of them reasoning) · 2.9 s · $0.0001 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Relace ·

  3. GLM 5.3 Flash

    Passed: Returned the right 3 rows.

    ```sql
    SELECT p.category,
           SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM orders o
    JOIN order_items oi ON oi.order_id = o.id
    JOIN products p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```

    508 tokens in, 98 out (20 of them reasoning) · 2.6 s · $0.0001 · 1 message on Pro · answered by z-ai/glm-5.3-flash via Wafer ·

  4. DeepSeek V4 Pro

    Passed: Returned the right 3 rows.

    ```sql
    SELECT p.category,
           SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM order_items oi
    JOIN products p ON p.id = oi.product_id
    JOIN orders o ON o.id = oi.order_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```

    506 tokens in, 187 out (102 of them reasoning) · 3.8 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Baidu ·

  5. GLM 5.3

    Passed: Returned the right 3 rows.

    ```sql
    SELECT p.category,
           SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM orders o
    JOIN order_items oi ON oi.order_id = o.id
    JOIN products p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```

    507 tokens in, 79 out · 1.2 s · $0.0002 · 1 message on Pro · answered by z-ai/glm-5.3 via Baidu ·

  6. Gemini 3.8 Flash

    Passed: Returned the right 3 rows.

    ```sql
    SELECT 
        p.category,
        SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM orders o
    JOIN order_items oi ON o.id = oi.order_id
    JOIN products p ON oi.product_id = p.id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```

    533 tokens in, 95 out · 3.0 s · $0.0008 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

  7. Claude Haiku 4.5

    Passed: Returned the right 3 rows.

    ```sql
    SELECT 
      p.category,
      SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM order_items oi
    JOIN products p ON oi.product_id = p.id
    JOIN orders o ON oi.order_id = o.id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```

    537 tokens in, 103 out · 1.2 s · $0.0011 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

  8. GPT-6 Sol

    Passed: Returned the right 3 rows.

    ```sql
    SELECT
        p.category,
        SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM orders AS o
    JOIN order_items AS oi ON oi.order_id = o.id
    JOIN products AS p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```

    493 tokens in, 84 out · 1.3 s · $0.0018 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  9. Kimi K3

    Passed: Returned the right 3 rows.

    ```sql
    SELECT p.category,
           SUM(oi.quantity * p.price_cents) / 100.0 AS revenue
    FROM order_items oi
    JOIN orders o ON o.id = oi.order_id
    JOIN products p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue DESC;
    ```

    580 tokens in, 90 out (6 of them reasoning) · 1.6 s · $0.0019 · 1 message on Pro · answered by moonshotai/kimi-k3 via Phala ·

  10. Claude Sonnet 5

    Passed: Returned the right 3 rows.

    ```sql
    SELECT p.category,
           SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM order_items oi
    JOIN products p ON p.id = oi.product_id
    JOIN orders o ON o.id = oi.order_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```

    820 tokens in, 141 out · 3.5 s · $0.0031 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

  11. Grok 4.7

    Passed: Returned the right 3 rows.

    ```sql
    SELECT p.category,
           SUM(oi.quantity * p.price_cents) / 100.0 AS revenue
    FROM order_items oi
    JOIN orders o ON o.id = oi.order_id
    JOIN products p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue DESC;
    ```

    1,739 tokens in, 411 out (330 of them reasoning) · 4.9 s · $0.0031 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  12. Claude Sonnet 5.5

    Passed: Returned the right 3 rows.

    ```sql
    SELECT p.category,
           SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM orders o
    JOIN order_items oi ON oi.order_id = o.id
    JOIN products p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```
    
    This assumes completed orders have `status = 'completed'`; adjust the string if your data uses a different value (e.g. `'Completed'`).

    824 tokens in, 187 out · 1.6 s · $0.0039 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·

  13. Gemini 3.1 Pro

    Passed: Returned the right 3 rows.

    ```sql
    SELECT 
        p.category, 
        SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM orders o
    JOIN order_items oi ON o.id = oi.order_id
    JOIN products p ON oi.product_id = p.id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```

    533 tokens in, 528 out (432 of them reasoning) · 7.4 s · $0.0074 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  14. GPT-6 Astra

    Passed: Returned the right 3 rows.

    ```sql
    SELECT
        p.category,
        SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM orders AS o
    JOIN order_items AS oi ON oi.order_id = o.id
    JOIN products AS p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```

    493 tokens in, 84 out · 2.2 s · $0.0091 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

  15. Claude Opus 5.5

    Passed: Returned the right 3 rows.

    ```sql
    SELECT
        p.category,
        ROUND(SUM(oi.quantity * p.price_cents) / 100.0, 2) AS revenue_dollars
    FROM orders o
    JOIN order_items oi ON oi.order_id = o.id
    JOIN products p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```
    
    This assumes completed orders have the status `'completed'`. If your data uses a different value, such as `'Completed'` or `'delivered'`, change the `WHERE` line to match.

    822 tokens in, 257 out (48 of them reasoning) · 4.3 s · $0.0092 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  16. Claude Fable 5.1

    Passed: Returned the right 3 rows.

    ```sql
    SELECT p.category,
           SUM(oi.quantity * p.price_cents) / 100.0 AS revenue_dollars
    FROM order_items oi
    JOIN orders o   ON o.id = oi.order_id
    JOIN products p ON p.id = oi.product_id
    WHERE o.status = 'completed'
    GROUP BY p.category
    ORDER BY revenue_dollars DESC;
    ```
    
    Assumes completed orders have `status = 'completed'`; adjust the literal if your data uses a different value.

    822 tokens in, 177 out · 4.5 s · $0.0190 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

More SQL prompts

The other SQL prompts, each with every model's reply, and the results across all five.

Questions

Which AI does best on “Revenue by category”?

All 16 models passed this SQL prompt's check (query result). The cheapest reply that passed was GPT-6 Luna's, at $0.000090, and the fastest too, in 1.2 s. The dearest reply, Claude Fable 5.1's, cost 210 times as much ($0.0190).

What does a reply to “Revenue by category” cost?

Through the models' APIs, what OpenRouter charged us ran from $0.000090 (GPT-6 Luna) to $0.0190 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.