Skip to content

Tested prompt · Agents and tool use

Book a meeting from a sentence: every AI model's reply, tested

We sent this everyday agents and tool use prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.

Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

15 of 16 models passed this agents and tool use prompt's check (tool calls); GLM 5.3 Flash didn't. The cheapest reply that passed was GPT-6 Luna's, at $0.000089; the fastest, Claude Haiku 4.5's in 1.2 s. The dearest reply, Claude Fable 5.1's, cost 147 times as much ($0.0131).

The prompt, as sent, and its check

Checked by tool calls, the same way for every model.

Book a meeting from a sentence (everyday)

You can call these tools:
- create_event(title: string, start: string in YYYY-MM-DDTHH:MM, duration_minutes: number, attendees: array of email addresses): adds an event to the user's calendar and invites the attendees.
- find_contact(name: string): looks up a contact's email address.

Today is Monday, 5 October 2026. The user's contacts include Priya Shah <priya@northwind.test>.
User: Put 30 minutes with Priya on my calendar this Thursday at 3pm to go over the Q4 plan.

Reply with only a JSON object, {"calls": [{"tool": "<name>", "arguments": {...}}]}, listing the tool calls to make now. If no call is right yet, reply {"calls": []}.

It must make exactly these calls, and no others:

[
  {
    "arguments": {
      "attendees": [
        "priya@northwind.test"
      ],
      "duration_minutes": 30,
      "start": "2026-10-08T15:00",
      "title": {
        "contains": "Q4"
      }
    },
    "tool": "create_event"
  }
]

Exactly what this prompt's replies are checked against, with every other prompt of our test runs.

Every model's result

All 16 models on this prompt, in catalog order.

Every model's reply to “Book a meeting from a sentence”
ModelResultCostTimeReply
Claude Fable 5.1AnthropicPassed: Made the 1 expected call.$0.01312.5 s87 tokens
Claude Opus 5.5AnthropicPassed: Made the 1 expected call.$0.00597.2 s87 tokens
Claude Sonnet 5.5AnthropicPassed: Made the 1 expected call.$0.00261.5 s89 tokens
Claude Sonnet 5AnthropicPassed: Made the 1 expected call.$0.00375.7 s86 tokens
Claude Haiku 4.5AnthropicPassed: Made the 1 expected call.$0.00101.2 s99 tokens
GPT-6 AstraOpenAIPassed: Made the 1 expected call.$0.00771.9 s55 tokens
GPT-6 SolOpenAIPassed: Made the 1 expected call.$0.00211.9 s57 tokens
GPT-6 LunaOpenAIPassed: Made the 1 expected call.$0.0000891.2 s53 tokens
Gemini 3.1 Pro (preview)GooglePassed: Made the 1 expected call.$0.00687.1 s110 tokens
Gemini 3.8 FlashGooglePassed: Made the 1 expected call.$0.00126.3 s70 tokens
DeepSeek V4.1 FlashDeepSeekPassed: Made the 1 expected call.$0.000231.4 s65 tokens
DeepSeek V4 ProDeepSeekPassed: Made the 1 expected call.$0.000263.5 s65 tokens
Grok 4.7xAIPassed: Made the 1 expected call.$0.00489.1 s59 tokens
Kimi K3MoonshotPassed: Made the 1 expected call.$0.00311.6 s75 tokens
GLM 5.3Z.aiPassed: Made the 1 expected call.$0.000561.7 s63 tokens
GLM 5.3 FlashZ.aiFailed: Expected create_event(attendees: ["priya@northwind.test"], duration_minutes: 30, start: "2026-10-08T15:00", title: {"contains":"Q4"}); got find_contact(name: "Priya Shah"); create_event(title: "Q4 plan review", start: "2026-10-08T15:00", duration_minutes: 30, attendees: ["priya@northwind.test"]).$0.000166.8 s79 tokens

Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.

Every reply

The reply that failed first, then the ones that passed, from the cheapest up.

  1. GLM 5.3 Flash

    Failed: Expected create_event(attendees: ["priya@northwind.test"], duration_minutes: 30, start: "2026-10-08T15:00", title: {"contains":"Q4"}); got find_contact(name: "Priya Shah"); create_event(title: "Q4 plan review", start: "2026-10-08T15:00", duration_minutes: 30, attendees: ["priya@northwind.test"]).

    {"calls": [{"tool": "find_contact", "arguments": {"name": "Priya Shah"}}, {"tool": "create_event", "arguments": {"title": "Q4 plan review", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    517 tokens in, 174 out (95 of them reasoning) · 6.8 s · $0.0002 · 1 message on Pro · answered by z-ai/glm-5.3-flash via Together ·

  2. GPT-6 Luna

    Passed: Made the 1 expected call.

    {"calls":[{"tool":"create_event","arguments":{"title":"Q4 plan","start":"2026-10-08T15:00","duration_minutes":30,"attendees":["priya@northwind.test"]}}]}

    499 tokens in, 78 out (25 of them reasoning) · 1.2 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  3. DeepSeek V4.1 Flash

    Passed: Made the 1 expected call.

    {"calls": [{"tool": "create_event", "arguments": {"title": "Q4 plan with Priya", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    552 tokens in, 152 out (87 of them reasoning) · 1.4 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Novita ·

  4. DeepSeek V4 Pro

    Passed: Made the 1 expected call.

    {"calls": [{"tool": "create_event", "arguments": {"title": "Q4 plan with Priya", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    524 tokens in, 148 out (83 of them reasoning) · 3.5 s · $0.0003 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via StreamLake ·

  5. GLM 5.3

    Passed: Made the 1 expected call.

    {"calls": [{"tool": "create_event", "arguments": {"title": "Q4 Plan Review with Priya", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    516 tokens in, 63 out · 1.7 s · $0.0006 · 1 message on Pro · answered by z-ai/glm-5.3 via Novita ·

  6. Claude Haiku 4.5

    Passed: Made the 1 expected call.

    ```json
    {
      "calls": [
        {
          "tool": "create_event",
          "arguments": {
            "title": "Q4 plan review",
            "start": "2026-10-08T15:00",
            "duration_minutes": 30,
            "attendees": ["priya@northwind.test"]
          }
        }
      ]
    }
    ```

    552 tokens in, 99 out · 1.2 s · $0.0010 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

  7. Gemini 3.8 Flash

    Passed: Made the 1 expected call.

    {"calls": [{"tool": "create_event", "arguments": {"title": "Go over the Q4 plan", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    545 tokens in, 211 out (141 of them reasoning) · 6.3 s · $0.0012 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

  8. GPT-6 Sol

    Passed: Made the 1 expected call.

    {"calls":[{"tool":"create_event","arguments":{"title":"Review Q4 plan with Priya","start":"2026-10-08T15:00","duration_minutes":30,"attendees":["priya@northwind.test"]}}]}

    499 tokens in, 112 out (55 of them reasoning) · 1.9 s · $0.0021 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  9. Claude Sonnet 5.5

    Passed: Made the 1 expected call.

    {"calls": [{"tool": "create_event", "arguments": {"title": "Q4 Plan Review with Priya", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    710 tokens in, 89 out · 1.5 s · $0.0026 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·

  10. Kimi K3

    Passed: Made the 1 expected call.

    {"calls": [{"tool": "create_event", "arguments": {"title": "Q4 plan review with Priya", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    588 tokens in, 110 out (35 of them reasoning) · 1.6 s · $0.0031 · 1 message on Pro · answered by moonshotai/kimi-k3 via Together ·

  11. Claude Sonnet 5

    Passed: Made the 1 expected call.

    {"calls": [{"tool": "create_event", "arguments": {"title": "Q4 Plan Review", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    706 tokens in, 231 out (145 of them reasoning) · 5.7 s · $0.0037 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

  12. Grok 4.7

    Passed: Made the 1 expected call.

    {"calls": [{"tool": "create_event", "arguments": {"title": "Q4 plan", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    1,742 tokens in, 773 out (714 of them reasoning) · 9.1 s · $0.0048 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  13. Claude Opus 5.5

    Passed: Made the 1 expected call.

    {"calls": [{"tool": "create_event", "arguments": {"title": "Q4 plan review with Priya", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    708 tokens in, 120 out (33 of them reasoning) · 7.2 s · $0.0059 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  14. Gemini 3.1 Pro

    Passed: Made the 1 expected call.

    ```json
    {
      "calls": [
        {
          "tool": "create_event",
          "arguments": {
            "title": "Go over the Q4 plan",
            "start": "2026-10-08T15:00",
            "duration_minutes": 30,
            "attendees": [
              "priya@northwind.test"
            ]
          }
        }
      ]
    }
    ```

    545 tokens in, 478 out (368 of them reasoning) · 7.1 s · $0.0068 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  15. GPT-6 Astra

    Passed: Made the 1 expected call.

    {"calls":[{"tool":"create_event","arguments":{"title":"Q4 plan review with Priya","start":"2026-10-08T15:00","duration_minutes":30,"attendees":["priya@northwind.test"]}}]}

    499 tokens in, 55 out · 1.9 s · $0.0077 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

  16. Claude Fable 5.1

    Passed: Made the 1 expected call.

    {"calls": [{"tool": "create_event", "arguments": {"title": "Q4 plan review with Priya", "start": "2026-10-08T15:00", "duration_minutes": 30, "attendees": ["priya@northwind.test"]}}]}

    708 tokens in, 87 out · 2.5 s · $0.0131 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

More agents and tool use prompts

The other agents and tool use prompts, each with every model's reply, and the results across all five.

Questions

Which AI does best on “Book a meeting from a sentence”?

15 of 16 models passed this agents and tool use prompt's check (tool calls); GLM 5.3 Flash didn't. The cheapest reply that passed was GPT-6 Luna's, at $0.000089; the fastest, Claude Haiku 4.5's in 1.2 s. The dearest reply, Claude Fable 5.1's, cost 147 times as much ($0.0131).

What does a reply to “Book a meeting from a sentence” cost?

Through the models' APIs, what OpenRouter charged us ran from $0.000089 (GPT-6 Luna) to $0.0131 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.