Tested prompt · Agents and tool use
Search, but don't book: every AI model's reply, tested
We sent this hard agents and tool use prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.
Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
All 16 models passed this agents and tool use prompt's check (tool calls). The cheapest reply that passed was GLM 5.3 Flash's, at $0.000079; the fastest, GLM 5.3's in 0.5 s. The dearest reply, Claude Fable 5.1's, cost 147 times as much ($0.0116).
The prompt, as sent, and its check
Checked by tool calls, the same way for every model.
Search, but don't book (hard)
You can call these tools: - search_flights(origin: string, destination: string, date: string in YYYY-MM-DD, max_stops: number): searches flights between airports given as IATA codes, like "LHR". - book_flight(flight_id: string): books a flight found by a search. Today is Monday, 5 October 2026. User: Find me nonstop flights from Boston to Denver on November 14. Don't book anything yet: I want to see the options first. Reply with only a JSON object, {"calls": [{"tool": "<name>", "arguments": {...}}]}, listing the tool calls to make now. If no call is right yet, reply {"calls": []}.
It must make exactly these calls, and no others:
[
{
"arguments": {
"date": "2026-11-14",
"destination": "DEN",
"max_stops": 0,
"origin": "BOS"
},
"tool": "search_flights"
}
]Exactly what this prompt's replies are checked against, with every other prompt of our test runs.
Every model's result
All 16 models on this prompt, in catalog order.
| Model | Result | Cost | Time | Reply |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | Passed: Made the 1 expected call. | $0.0116 | 5.0 s | 66 tokens |
| Claude Opus 5.5Anthropic | Passed: Made the 1 expected call. | $0.0046 | 3.1 s | 66 tokens |
| Claude Sonnet 5.5Anthropic | Passed: Made the 1 expected call. | $0.0023 | 1.4 s | 66 tokens |
| Claude Sonnet 5Anthropic | Passed: Made the 1 expected call. | $0.0020 | 1.9 s | 66 tokens |
| Claude Haiku 4.5Anthropic | Passed: Made the 1 expected call. | $0.00096 | 1.1 s | 85 tokens |
| GPT-6 AstraOpenAI | Passed: Made the 1 expected call. | $0.0069 | 1.5 s | 41 tokens |
| GPT-6 SolOpenAI | Passed: Made the 1 expected call. | $0.0014 | 1.1 s | 41 tokens |
| GPT-6 LunaOpenAI | Passed: Made the 1 expected call. | $0.000085 | 2.2 s | 43 tokens |
| Gemini 3.1 Pro (preview)Google | Passed: Made the 1 expected call. | $0.0061 | 4.7 s | 85 tokens |
| Gemini 3.8 FlashGoogle | Passed: Made the 1 expected call. | $0.00089 | 6.3 s | 51 tokens |
| DeepSeek V4.1 FlashDeepSeek | Passed: Made the 1 expected call. | $0.00014 | 1.2 s | 50 tokens |
| DeepSeek V4 ProDeepSeek | Passed: Made the 1 expected call. | $0.00061 | 1.3 s | 50 tokens |
| Grok 4.7xAI | Passed: Made the 1 expected call. | $0.0020 | 2.5 s | 48 tokens |
| Kimi K3Moonshot | Passed: Made the 1 expected call. | $0.0034 | 2.0 s | 48 tokens |
| GLM 5.3Z.ai | Passed: Made the 1 expected call. | $0.00054 | 0.5 s | 37 tokens |
| GLM 5.3 FlashZ.ai | Passed: Made the 1 expected call. | $0.000079 | 2.3 s | 49 tokens |
Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.
Every reply
Every reply passed: here they are from the cheapest up.
GLM 5.3 Flash
Passed: Made the 1 expected call.
{"calls": [{"tool": "search_flights", "arguments": {"origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0}}]}504 tokens in, 68 out (19 of them reasoning) · 2.3 s · $0.0001 · 1 message on Pro · answered by z-ai/glm-5.3-flash via AtlasCloud ·
GPT-6 Luna
Passed: Made the 1 expected call.
{"calls":[{"tool":"search_flights","arguments":{"origin":"BOS","destination":"DEN","date":"2026-11-14","max_stops":0}}]}487 tokens in, 72 out (29 of them reasoning) · 2.2 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·
DeepSeek V4.1 Flash
Passed: Made the 1 expected call.
{"calls": [{"tool": "search_flights", "arguments": {"origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0}}]}537 tokens in, 75 out (25 of them reasoning) · 1.2 s · $0.0001 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via StreamLake ·
GLM 5.3
Passed: Made the 1 expected call.
{"calls":[{"tool":"search_flights","arguments":{"origin":"BOS","destination":"DEN","date":"2026-11-14","max_stops":0}}]}527 tokens in, 38 out (1 of them reasoning) · 0.5 s · $0.0005 · 1 message on Pro · answered by z-ai/glm-5.3 via Wafer ·
DeepSeek V4 Pro
Passed: Made the 1 expected call.
{"calls": [{"tool": "search_flights", "arguments": {"origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0}}]}509 tokens in, 140 out (90 of them reasoning) · 1.3 s · $0.0006 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Wafer ·
Gemini 3.8 Flash
Passed: Made the 1 expected call.
{"calls": [{"tool": "search_flights", "arguments": {"origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0}}]}528 tokens in, 133 out (82 of them reasoning) · 6.3 s · $0.0009 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·
Claude Haiku 4.5
Passed: Made the 1 expected call.
```json { "calls": [ { "tool": "search_flights", "arguments": { "origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0 } } ] } ```531 tokens in, 85 out · 1.1 s · $0.0010 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·
GPT-6 Sol
Passed: Made the 1 expected call.
{"calls":[{"tool":"search_flights","arguments":{"origin":"BOS","destination":"DEN","date":"2026-11-14","max_stops":0}}]}487 tokens in, 41 out · 1.1 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Claude Sonnet 5
Passed: Made the 1 expected call.
{"calls": [{"tool": "search_flights", "arguments": {"origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0}}]}675 tokens in, 66 out · 1.9 s · $0.0020 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·
Grok 4.7
Passed: Made the 1 expected call.
{"calls": [{"tool": "search_flights", "arguments": {"origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0}}]}1,729 tokens in, 195 out (147 of them reasoning) · 2.5 s · $0.0020 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
Claude Sonnet 5.5
Passed: Made the 1 expected call.
{"calls": [{"tool": "search_flights", "arguments": {"origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0}}]}679 tokens in, 66 out · 1.4 s · $0.0023 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·
Kimi K3
Passed: Made the 1 expected call.
{"calls":[{"tool":"search_flights","arguments":{"origin":"BOS","destination":"DEN","date":"2026-11-14","max_stops":0}}]}576 tokens in, 133 out (85 of them reasoning) · 2.0 s · $0.0034 · 1 message on Pro · answered by moonshotai/kimi-k3 via Together ·
Claude Opus 5.5
Passed: Made the 1 expected call.
{"calls": [{"tool": "search_flights", "arguments": {"origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0}}]}677 tokens in, 66 out · 3.1 s · $0.0046 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·
Gemini 3.1 Pro
Passed: Made the 1 expected call.
```json { "calls": [ { "tool": "search_flights", "arguments": { "origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0 } } ] } ```528 tokens in, 418 out (333 of them reasoning) · 4.7 s · $0.0061 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·
GPT-6 Astra
Passed: Made the 1 expected call.
{"calls":[{"tool":"search_flights","arguments":{"origin":"BOS","destination":"DEN","date":"2026-11-14","max_stops":0}}]}487 tokens in, 41 out · 1.5 s · $0.0069 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·
Claude Fable 5.1
Passed: Made the 1 expected call.
{"calls": [{"tool": "search_flights", "arguments": {"origin": "BOS", "destination": "DEN", "date": "2026-11-14", "max_stops": 0}}]}677 tokens in, 66 out · 5.0 s · $0.0116 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·
More agents and tool use prompts
The other agents and tool use prompts, each with every model's reply, and the results across all five.
Questions
Which AI does best on “Search, but don't book”?
All 16 models passed this agents and tool use prompt's check (tool calls). The cheapest reply that passed was GLM 5.3 Flash's, at $0.000079; the fastest, GLM 5.3's in 0.5 s. The dearest reply, Claude Fable 5.1's, cost 147 times as much ($0.0116).
What does a reply to “Search, but don't book” cost?
Through the models' APIs, what OpenRouter charged us ran from $0.000079 (GLM 5.3 Flash) to $0.0116 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.