Tested prompt · RAG and answering from documents
A question the handbook doesn't answer: every AI model's reply, tested
We sent this hard RAG and answering from documents prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.
Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
All 16 models passed this RAG and answering from documents prompt's check (key facts). The cheapest reply that passed was GPT-6 Luna's, at $0.000071; the fastest, GLM 5.3's in 0.6 s. The dearest reply, Claude Fable 5.1's, cost 226 times as much ($0.0160).
The prompt, as sent, and its check
Checked by key facts, the same way for every model.
A question the handbook doesn't answer (hard)
[9 lines every RAG and answering from documents prompt of ours shares, word for word: the whole prompt, on the methods page] Question: How many sick days do employees get a year?
The handbook doesn't answer this: the reply must say so, and not make up a number.
Exactly what this prompt's replies are checked against, with every other prompt of our test runs.
Every model's result
All 16 models on this prompt, in catalog order.
| Model | Result | Cost | Time | Reply |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | Passed: Said the handbook doesn't answer it. | $0.0160 | 4.0 s | 100 tokens |
| Claude Opus 5.5Anthropic | Passed: Said the handbook doesn't answer it. | $0.0057 | 3.4 s | 50 tokens |
| Claude Sonnet 5.5Anthropic | Passed: Said the handbook doesn't answer it. | $0.0028 | 1.4 s | 60 tokens |
| Claude Sonnet 5Anthropic | Passed: Said the handbook doesn't answer it. | $0.0019 | 1.8 s | 12 tokens |
| Claude Haiku 4.5Anthropic | Passed: Said the handbook doesn't answer it. | $0.00092 | 1.2 s | 42 tokens |
| GPT-6 AstraOpenAI | Passed: Said the handbook doesn't answer it. | $0.0073 | 1.4 s | 19 tokens |
| GPT-6 SolOpenAI | Passed: Said the handbook doesn't answer it. | $0.0014 | 0.9 s | 15 tokens |
| GPT-6 LunaOpenAI | Passed: Said the handbook doesn't answer it. | $0.000071 | 1.9 s | 14 tokens |
| Gemini 3.1 Pro (preview)Google | Passed: Said the handbook doesn't answer it. | $0.0046 | 5.6 s | 8 tokens |
| Gemini 3.8 FlashGoogle | Passed: Said the handbook doesn't answer it. | $0.00055 | 2.9 s | 8 tokens |
| DeepSeek V4.1 FlashDeepSeek | Passed: Said the handbook doesn't answer it. | $0.00013 | 1.1 s | 22 tokens |
| DeepSeek V4 ProDeepSeek | Passed: Said the handbook doesn't answer it. | $0.00023 | 1.9 s | 7 tokens |
| Grok 4.7xAI | Passed: Said the handbook doesn't answer it. | $0.0021 | 1.8 s | 6 tokens |
| Kimi K3Moonshot | Passed: Said the handbook doesn't answer it. | $0.0010 | 0.9 s | 22 tokens |
| GLM 5.3Z.ai | Passed: Said the handbook doesn't answer it. | $0.00046 | 0.6 s | 37 tokens |
| GLM 5.3 FlashZ.ai | Passed: Said the handbook doesn't answer it. | $0.000098 | 2.9 s | 58 tokens |
Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.
Every reply
Every reply passed: here they are from the cheapest up.
GPT-6 Luna
Passed: Said the handbook doesn't answer it.
Sick days aren’t specified in the handbook.
637 tokens in, 14 out · 1.9 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·
GLM 5.3 Flash
Passed: Said the handbook doesn't answer it.
It isn't in the handbook. The excerpt covers working hours (§1), remote work (§2), equipment (§3), parental leave (§4), holidays (§5), learning budget (§6), and amendments (§7), but says nothing about sick days.
657 tokens in, 92 out (34 of them reasoning) · 2.9 s · $0.0001 · 1 message on Pro · answered by z-ai/glm-5.3-flash via Wafer ·
DeepSeek V4.1 Flash
Passed: Said the handbook doesn't answer it.
That isn't in the handbook — none of the sections (§1–§7) cover sick days.
684 tokens in, 33 out (11 of them reasoning) · 1.1 s · $0.0001 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via StreamLake ·
DeepSeek V4 Pro
Passed: Said the handbook doesn't answer it.
The handbook doesn't say.
656 tokens in, 74 out (67 of them reasoning) · 1.9 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Ionstream ·
GLM 5.3
Passed: Said the handbook doesn't answer it.
That isn't in the handbook. The excerpt covers working hours, remote work, equipment, parental leave, holidays, learning, and amendments — but it says nothing about sick days.
680 tokens in, 38 out (1 of them reasoning) · 0.6 s · $0.0005 · 1 message on Pro · answered by z-ai/glm-5.3 via Wafer ·
Gemini 3.8 Flash
Passed: Said the handbook doesn't answer it.
It isn't in the handbook.
696 tokens in, 8 out · 2.9 s · $0.0006 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·
Claude Haiku 4.5
Passed: Said the handbook doesn't answer it.
That isn't in the handbook. The handbook covers working hours, remote work, equipment, parental leave, holidays, and learning budget, but doesn't specify an annual allocation of sick days.
707 tokens in, 42 out · 1.2 s · $0.0009 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·
Kimi K3
Passed: Said the handbook doesn't answer it.
The handbook doesn't say — there's no section covering sick days.
727 tokens in, 36 out (14 of them reasoning) · 0.9 s · $0.0010 · 1 message on Pro · answered by moonshotai/kimi-k3 via Wafer ·
GPT-6 Sol
Passed: Said the handbook doesn't answer it.
The number of sick days isn’t in the handbook.
637 tokens in, 15 out · 0.9 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Claude Sonnet 5
Passed: Said the handbook doesn't answer it.
That isn't in the handbook.
889 tokens in, 12 out · 1.8 s · $0.0019 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·
Grok 4.7
Passed: Said the handbook doesn't answer it.
It isn't in the handbook.
1,871 tokens in, 96 out (90 of them reasoning) · 1.8 s · $0.0021 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
Claude Sonnet 5.5
Passed: Said the handbook doesn't answer it.
That isn't in the handbook. The excerpt covers working hours, remote work, equipment, parental leave, holidays, learning budget, and amendments, but says nothing about sick days.
893 tokens in, 60 out · 1.4 s · $0.0028 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·
Gemini 3.1 Pro
Passed: Said the handbook doesn't answer it.
That isn't in the handbook.
696 tokens in, 266 out (258 of them reasoning) · 5.6 s · $0.0046 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·
Claude Opus 5.5
Passed: Said the handbook doesn't answer it.
That isn't in the handbook. The excerpt doesn't mention sick days; the closest sections cover holidays (§5) and parental leave (§4).
891 tokens in, 66 out (16 of them reasoning) · 3.4 s · $0.0057 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·
GPT-6 Astra
Passed: Said the handbook doesn't answer it.
The number of sick days employees get per year isn’t in the handbook.
637 tokens in, 19 out · 1.4 s · $0.0073 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·
Claude Fable 5.1
Passed: Said the handbook doesn't answer it.
The handbook excerpt doesn't say anything about sick days. It covers working hours (§1), remote work (§2), equipment (§3), parental leave (§4), holidays (§5), learning budget (§6) and amendments (§7), but there is no sick leave section, so I can't answer this from the handbook.
891 tokens in, 100 out · 4.0 s · $0.0160 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·
More RAG and answering from documents prompts
The other RAG and answering from documents prompts, each with every model's reply, and the results across all five.
Questions
Which AI does best on “A question the handbook doesn't answer”?
All 16 models passed this RAG and answering from documents prompt's check (key facts). The cheapest reply that passed was GPT-6 Luna's, at $0.000071; the fastest, GLM 5.3's in 0.6 s. The dearest reply, Claude Fable 5.1's, cost 226 times as much ($0.0160).
What does a reply to “A question the handbook doesn't answer” cost?
Through the models' APIs, what OpenRouter charged us ran from $0.000071 (GPT-6 Luna) to $0.0160 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.