Skip to content

Tested prompt · Agents and tool use

Two calls with a unit conversion: every AI model's reply, tested

We sent this hard agents and tool use prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.

Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

All 16 models passed this agents and tool use prompt's check (tool calls). The cheapest reply that passed was GPT-6 Luna's, at $0.000092; the fastest, GLM 5.3's in 0.7 s. The dearest reply, Claude Fable 5.1's, cost 157 times as much ($0.0145).

The prompt, as sent, and its check

Checked by tool calls, the same way for every model.

Two calls with a unit conversion (hard)

You can call these tools:
- set_thermostat(room: "living room" | "bedroom" | "office", celsius: number to one decimal place): sets a room's target temperature.
- get_temperature(room: "living room" | "bedroom" | "office"): reads a room's current temperature.

User: Set the living room to 72°F, and the bedroom two degrees Celsius cooler than that.

Reply with only a JSON object, {"calls": [{"tool": "<name>", "arguments": {...}}]}, listing the tool calls to make now. If no call is right yet, reply {"calls": []}.

It must make exactly these calls, and no others:

[
  {
    "arguments": {
      "celsius": 22.2,
      "room": "living room"
    },
    "tool": "set_thermostat"
  },
  {
    "arguments": {
      "celsius": 20.2,
      "room": "bedroom"
    },
    "tool": "set_thermostat"
  }
]

Exactly what this prompt's replies are checked against, with every other prompt of our test runs.

Every model's result

All 16 models on this prompt, in catalog order.

Every model's reply to “Two calls with a unit conversion”
ModelResultCostTimeReply
Claude Fable 5.1AnthropicPassed: Made the 2 expected calls.$0.01454.8 s129 tokens
Claude Opus 5.5AnthropicPassed: Made the 2 expected calls.$0.00643.4 s87 tokens
Claude Sonnet 5.5AnthropicPassed: Made the 2 expected calls.$0.00383.7 s87 tokens
Claude Sonnet 5AnthropicPassed: Made the 2 expected calls.$0.00283.8 s87 tokens
Claude Haiku 4.5AnthropicPassed: Made the 2 expected calls.$0.00111.2 s109 tokens
GPT-6 AstraOpenAIPassed: Made the 2 expected calls.$0.00742.3 s54 tokens
GPT-6 SolOpenAIPassed: Made the 2 expected calls.$0.00182.8 s56 tokens
GPT-6 LunaOpenAIPassed: Made the 2 expected calls.$0.0000921.9 s56 tokens
Gemini 3.1 Pro (preview)GooglePassed: Made the 2 expected calls.$0.00636.9 s103 tokens
Gemini 3.8 FlashGooglePassed: Made the 2 expected calls.$0.000983.9 s61 tokens
DeepSeek V4.1 FlashDeepSeekPassed: Made the 2 expected calls.$0.000242.0 s65 tokens
DeepSeek V4 ProDeepSeekPassed: Made the 2 expected calls.$0.000193.3 s66 tokens
Grok 4.7xAIPassed: Made the 2 expected calls.$0.00294.5 s61 tokens
Kimi K3MoonshotPassed: Made the 2 expected calls.$0.00121.7 s72 tokens
GLM 5.3Z.aiPassed: Made the 2 expected calls.$0.000760.7 s61 tokens
GLM 5.3 FlashZ.aiPassed: Made the 2 expected calls.$0.000203.4 s61 tokens

Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.

Every reply

Every reply passed: here they are from the cheapest up.

  1. GPT-6 Luna

    Passed: Made the 2 expected calls.

    {"calls":[{"tool":"set_thermostat","arguments":{"room":"living room","celsius":22.2}},{"tool":"set_thermostat","arguments":{"room":"bedroom","celsius":20.2}}]}

    472 tokens in, 90 out (34 of them reasoning) · 1.9 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  2. DeepSeek V4 Pro

    Passed: Made the 2 expected calls.

    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    492 tokens in, 174 out (108 of them reasoning) · 3.3 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Baidu ·

  3. GLM 5.3 Flash

    Passed: Made the 2 expected calls.

    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    487 tokens in, 180 out (119 of them reasoning) · 3.4 s · $0.0002 · 1 message on Pro · answered by z-ai/glm-5.3-flash via Wafer ·

  4. DeepSeek V4.1 Flash

    Passed: Made the 2 expected calls.

    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    520 tokens in, 236 out (171 of them reasoning) · 2.0 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via StreamLake ·

  5. GLM 5.3

    Passed: Made the 2 expected calls.

    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    510 tokens in, 93 out (32 of them reasoning) · 0.7 s · $0.0008 · 1 message on Pro · answered by z-ai/glm-5.3 via Wafer ·

  6. Gemini 3.8 Flash

    Passed: Made the 2 expected calls.

    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    504 tokens in, 160 out (99 of them reasoning) · 3.9 s · $0.0010 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

  7. Claude Haiku 4.5

    Passed: Made the 2 expected calls.

    ```json
    {
      "calls": [
        {
          "tool": "set_thermostat",
          "arguments": {
            "room": "living room",
            "celsius": 22.2
          }
        },
        {
          "tool": "set_thermostat",
          "arguments": {
            "room": "bedroom",
            "celsius": 20.2
          }
        }
      ]
    }
    ```

    508 tokens in, 109 out · 1.2 s · $0.0011 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

  8. Kimi K3

    Passed: Made the 2 expected calls.

    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    559 tokens in, 101 out (29 of them reasoning) · 1.7 s · $0.0012 · 1 message on Pro · answered by moonshotai/kimi-k3 via Wafer ·

  9. GPT-6 Sol

    Passed: Made the 2 expected calls.

    {"calls":[{"tool":"set_thermostat","arguments":{"room":"living room","celsius":22.2}},{"tool":"set_thermostat","arguments":{"room":"bedroom","celsius":20.2}}]}

    472 tokens in, 88 out (32 of them reasoning) · 2.8 s · $0.0018 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  10. Claude Sonnet 5

    Passed: Made the 2 expected calls.

    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    652 tokens in, 154 out (67 of them reasoning) · 3.8 s · $0.0028 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

  11. Grok 4.7

    Passed: Made the 2 expected calls.

    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    1,709 tokens in, 392 out (331 of them reasoning) · 4.5 s · $0.0029 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  12. Claude Sonnet 5.5

    Passed: Made the 2 expected calls.

    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    656 tokens in, 215 out (128 of them reasoning) · 3.7 s · $0.0038 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·

  13. Gemini 3.1 Pro

    Passed: Made the 2 expected calls.

    {
      "calls": [
        {
          "tool": "set_thermostat",
          "arguments": {
            "room": "living room",
            "celsius": 22.2
          }
        },
        {
          "tool": "set_thermostat",
          "arguments": {
            "room": "bedroom",
            "celsius": 20.2
          }
        }
      ]
    }

    504 tokens in, 441 out (338 of them reasoning) · 6.9 s · $0.0063 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  14. Claude Opus 5.5

    Passed: Made the 2 expected calls.

    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    654 tokens in, 161 out (74 of them reasoning) · 3.4 s · $0.0064 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  15. GPT-6 Astra

    Passed: Made the 2 expected calls.

    {"calls":[{"tool":"set_thermostat","arguments":{"room":"living room","celsius":22.2}},{"tool":"set_thermostat","arguments":{"room":"bedroom","celsius":20.2}}]}

    472 tokens in, 54 out · 2.3 s · $0.0074 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

  16. Claude Fable 5.1

    Passed: Made the 2 expected calls.

    72°F = (72-32)×5/9 = 22.222... → 22.2°C. Bedroom = 20.2°C.
    
    {"calls": [{"tool": "set_thermostat", "arguments": {"room": "living room", "celsius": 22.2}}, {"tool": "set_thermostat", "arguments": {"room": "bedroom", "celsius": 20.2}}]}

    654 tokens in, 129 out · 4.8 s · $0.0145 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

More agents and tool use prompts

The other agents and tool use prompts, each with every model's reply, and the results across all five.

Questions

Which AI does best on “Two calls with a unit conversion”?

All 16 models passed this agents and tool use prompt's check (tool calls). The cheapest reply that passed was GPT-6 Luna's, at $0.000092; the fastest, GLM 5.3's in 0.7 s. The dearest reply, Claude Fable 5.1's, cost 157 times as much ($0.0145).

What does a reply to “Two calls with a unit conversion” cost?

Through the models' APIs, what OpenRouter charged us ran from $0.000092 (GPT-6 Luna) to $0.0145 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.