Skip to content

Tested prompt · RAG and answering from documents

Core hours and start times: every AI model's reply, tested

We sent this everyday RAG and answering from documents prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.

Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

All 16 models passed this RAG and answering from documents prompt's check (key facts). The cheapest reply that passed was GPT-6 Luna's, at $0.000079; the fastest, GLM 5.3's in 0.5 s. The dearest reply, Claude Fable 5.1's, cost 169 times as much ($0.0135).

The prompt, as sent, and its check

Checked by key facts, the same way for every model.

Core hours and start times (everyday)

[9 lines every RAG and answering from documents prompt of ours shares, word for word: the whole prompt, on the methods page]
Question: What are the core hours, and how early can I start?

The answer must state “10:00”, “16:00”, “07:30”, “§1”.

Exactly what this prompt's replies are checked against, with every other prompt of our test runs.

Every model's result

All 16 models on this prompt, in catalog order.

Every model's reply to “Core hours and start times”
ModelResultCostTimeReply
Claude Fable 5.1AnthropicPassed: Stated all 4 facts with the section.$0.01353.7 s49 tokens
Claude Opus 5.5AnthropicPassed: Stated all 4 facts with the section.$0.00542.6 s49 tokens
Claude Sonnet 5.5AnthropicPassed: Stated all 4 facts with the section.$0.00271.4 s52 tokens
Claude Sonnet 5AnthropicPassed: Stated all 4 facts with the section.$0.00222.1 s40 tokens
Claude Haiku 4.5AnthropicPassed: Stated all 4 facts with the section.$0.000911.1 s40 tokens
GPT-6 AstraOpenAIPassed: Stated all 4 facts with the section.$0.00802.4 s31 tokens
GPT-6 SolOpenAIPassed: Stated all 4 facts with the section.$0.00161.1 s30 tokens
GPT-6 LunaOpenAIPassed: Stated all 4 facts with the section.$0.0000791.0 s31 tokens
Gemini 3.1 Pro (preview)GooglePassed: Stated all 4 facts with the section.$0.00525.8 s33 tokens
Gemini 3.8 FlashGooglePassed: Stated all 4 facts with the section.$0.000652.9 s33 tokens
DeepSeek V4.1 FlashDeepSeekPassed: Stated all 4 facts with the section.$0.000151.0 s28 tokens
DeepSeek V4 ProDeepSeekPassed: Stated all 4 facts with the section.$0.000572.5 s28 tokens
Grok 4.7xAIPassed: Stated all 4 facts with the section.$0.00221.9 s32 tokens
Kimi K3MoonshotPassed: Stated all 4 facts with the section.$0.00111.2 s44 tokens
GLM 5.3Z.aiPassed: Stated all 4 facts with the section.$0.000720.5 s29 tokens
GLM 5.3 FlashZ.aiPassed: Stated all 4 facts with the section.$0.000223.9 s29 tokens

Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.

Every reply

Every reply passed: here they are from the cheapest up.

  1. GPT-6 Luna

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00–16:00, and you can start as early as 07:30. (§1)

    640 tokens in, 31 out · 1.0 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  2. DeepSeek V4.1 Flash

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start as early as 07:30 (§1).

    687 tokens in, 49 out (21 of them reasoning) · 1.0 s · $0.0001 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via StreamLake ·

  3. GLM 5.3 Flash

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start as early as 07:30 (§1).

    660 tokens in, 84 out (55 of them reasoning) · 3.9 s · $0.0002 · 1 message on Pro · answered by z-ai/glm-5.3-flash via Wafer ·

  4. DeepSeek V4 Pro

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start as early as 07:30 (§1).

    659 tokens in, 109 out (81 of them reasoning) · 2.5 s · $0.0006 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Alibaba ·

  5. Gemini 3.8 Flash

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start as early as 07:30 (§1).

    699 tokens in, 33 out · 2.9 s · $0.0006 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

  6. GLM 5.3

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start any time from 07:30 (§1).

    683 tokens in, 30 out (1 of them reasoning) · 0.5 s · $0.0007 · 1 message on Pro · answered by z-ai/glm-5.3 via Wafer ·

  7. Claude Haiku 4.5

    Passed: Stated all 4 facts with the section.

    According to §1, core hours are 10:00 to 16:00, and you can start any time between 07:30 and 10:00.

    710 tokens in, 40 out · 1.1 s · $0.0009 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

  8. Kimi K3

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start any time between 07:30 and 10:00. (§1)

    730 tokens in, 67 out (23 of them reasoning) · 1.2 s · $0.0011 · 1 message on Pro · answered by moonshotai/kimi-k3 via Wafer ·

  9. GPT-6 Sol

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00–16:00, and you can start as early as 07:30 (§1).

    640 tokens in, 30 out · 1.1 s · $0.0016 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  10. Claude Sonnet 5

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start any time between 07:30 and 10:00 (§1).

    891 tokens in, 40 out · 2.1 s · $0.0022 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

  11. Grok 4.7

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start any time between 07:30 and 10:00 (§1).

    1,874 tokens in, 125 out (93 of them reasoning) · 1.9 s · $0.0022 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  12. Claude Sonnet 5.5

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start any time between 07:30 and 10:00, so the earliest start is 07:30. (§1)

    895 tokens in, 52 out · 1.4 s · $0.0027 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·

  13. Gemini 3.1 Pro

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start as early as 07:30 (§1).

    699 tokens in, 316 out (283 of them reasoning) · 5.8 s · $0.0052 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  14. Claude Opus 5.5

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start as early as 07:30 (any time between 07:30 and 10:00) (§1).

    893 tokens in, 49 out · 2.6 s · $0.0054 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  15. GPT-6 Astra

    Passed: Stated all 4 facts with the section.

    Core hours are **10:00–16:00**, and you can start as early as **07:30** (§1).

    640 tokens in, 31 out · 2.4 s · $0.0080 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

  16. Claude Fable 5.1

    Passed: Stated all 4 facts with the section.

    Core hours are 10:00 to 16:00, and you can start as early as 07:30 (any time between 07:30 and 10:00) (§1).

    893 tokens in, 49 out · 3.7 s · $0.0135 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

More RAG and answering from documents prompts

The other RAG and answering from documents prompts, each with every model's reply, and the results across all five.

Questions

Which AI does best on “Core hours and start times”?

All 16 models passed this RAG and answering from documents prompt's check (key facts). The cheapest reply that passed was GPT-6 Luna's, at $0.000079; the fastest, GLM 5.3's in 0.5 s. The dearest reply, Claude Fable 5.1's, cost 169 times as much ($0.0135).

What does a reply to “Core hours and start times” cost?

Through the models' APIs, what OpenRouter charged us ran from $0.000079 (GPT-6 Luna) to $0.0135 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.