Skip to content

Best AI · Research

The best AI for research, from our test runs

On our research prompts (answering from documents, summarizing and analyzing data), DeepSeek V4.1 Flash, Grok 4.7, Gemini 3.1 Pro, Claude Opus 5.5, and GPT-6 Astra each passed 15 of 15, the most; DeepSeek V4.1 Flash is first on the tie-break (the smaller message, then the lower cost). Every model's result, what each costs, and what the test doesn't cover.

Test runs and prices checked . Updated .

Short answer

DeepSeek V4.1 Flash, Grok 4.7, Gemini 3.1 Pro, Claude Opus 5.5, and GPT-6 Astra each passed 15 of 15, the most; DeepSeek V4.1 Flash is first on the tie-break (the smaller message, then the lower cost). This is our own test, run on September 27, 2026: every prompt, reply and score is published.

Ranked by our test runs

This is our own test: the same prompts sent to every model through llmwise, each reply checked the same way, with every prompt, reply and score published. The research prompts test RAG and answering from documents, summarization, and data analysis over material we give the model; they don't test web search.

Models ranked on RAG and answering from documents, summarization, and data analysis
#ModelPassedHard onesFree trialOn ProCost per reply
1DeepSeek V4.1 FlashDeepSeek15 of 156 of 6Yes60 a day$0.0003
2Grok 4.7xAI15 of 156 of 6Yes250 a month$0.0047
3Gemini 3.1 Pro (preview)Google15 of 156 of 6Yes125 a month$0.0100
4Claude Opus 5.5Anthropic15 of 156 of 6No62 a month$0.0100
5GPT-6 AstraOpenAI15 of 156 of 6No31 a month$0.0114
6GPT-6 LunaOpenAI14 of 156 of 6Yes60 a day$0.0001
7DeepSeek V4 ProDeepSeek14 of 156 of 6Yes250 a month$0.0010
8GLM 5.3Z.ai14 of 156 of 6Yes250 a month$0.0010
9GPT-6 SolOpenAI14 of 156 of 6Yes125 a month$0.0024
10Gemini 3.8 FlashGoogle14 of 155 of 6Yes250 a month$0.0014
11Kimi K3Moonshot13 of 156 of 6Yes125 a month$0.0039
12Claude Fable 5.1Anthropic13 of 156 of 6No31 a month$0.0196
13Claude Haiku 4.5Anthropic13 of 155 of 6Yes250 a month$0.0015
14GLM 5.3 FlashZ.ai12 of 155 of 6Yes60 a day$0.0002
15Claude Sonnet 5Anthropic12 of 155 of 6Yes125 a month$0.0040
15 prompts per model (RAG and answering from documents, summarization, and data analysis), run on September 27, 2026. Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand. Cost per reply is what OpenRouter charged us on average; in llmwise you pay per message.

The prompts, every reply and how each was scored: our test runs.

The best pick at each price

The same results, by what a message counts as on Pro: the best model at each size of message, then the rest of that size.

  • 60 a day on Pro (everyday models)

    DeepSeek V4.1 Flash, 15 of 15 passed; then GPT-6 Luna 14 of 15 and GLM 5.3 Flash 12 of 15.

  • 250 a month on Pro

    Grok 4.7, 15 of 15 passed; then DeepSeek V4 Pro 14 of 15, GLM 5.3 14 of 15, Gemini 3.8 Flash 14 of 15, and Claude Haiku 4.5 13 of 15.

  • 125 a month on Pro

    Gemini 3.1 Pro, 15 of 15 passed; then GPT-6 Sol 14 of 15, Kimi K3 13 of 15, and Claude Sonnet 5 12 of 15.

  • 62 a month on Pro

    Claude Opus 5.5, 15 of 15 passed.

  • 31 a month on Pro

    GPT-6 Astra, 15 of 15 passed; then Claude Fable 5.1 13 of 15.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Where the top models differ

The top 5, job by job.

Prompts passed on each job, by model
JobDeepSeek V4.1 FlashGrok 4.7Gemini 3.1 ProClaude Opus 5.5GPT-6 Astra
RAG and answering from documents5 of 55 of 55 of 55 of 55 of 5
Summarization5 of 55 of 55 of 55 of 55 of 5
Data analysis5 of 55 of 55 of 55 of 55 of 5

One prompt, side by side

A everyday summarization prompt, “An article in three bullets”, where the models didn't all pass.

The prompt: An article in three bullets (everyday)
Summarize this article in exactly 3 bullet points, at most 60 words in total.

The city of Aldmoor ended its six-month trial of protected bike lanes on Market Street last week, and the council will vote in November on whether to keep them. Bike trips on the street rose from about 900 to 2,300 a day during the trial, according to automatic counters. Car traffic fell by 14%, but travel times for drivers rose by just over a minute at rush hour. Shop owners are split: a survey of 60 businesses found 27 reported more customers, 19 reported fewer, and the rest saw no change. Two cyclist injuries were recorded on the street during the trial, down from nine in the same months last year. The lanes cost €410,000 to install; making them permanent would cost another €1.1 million for kerbs and new traffic lights.
  • Claude Fable 5.1

    Failed: Graded 4.3 of 5 on average (lowest 4); but 62 words, over the 60 allowed.

    - Aldmoor's six-month protected bike lane trial on Market Street ended; council votes in November on making it permanent.
    - Daily bike trips rose from 900 to 2,300, car traffic fell 14%, rush-hour driver times rose about a minute, and cyclist injuries dropped from nine to two.
    - Businesses were split (27 gained customers, 19 lost); permanence would cost €1.1 million beyond the €410,000 installation.

    727 tokens in, 150 out · 5.1 s · $0.0164 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

  • Gemini 3.1 Pro

    Passed: Graded 4.3 of 5 on average (lowest 4).

    * Daily bike trips jumped to 2,300 and cyclist injuries dropped sharply.
    * Car traffic fell 14% with minor delays, while local business impacts were mixed.
    * The city council votes in November on spending €1.1 million to make the lanes permanent.

    560 tokens in, 636 out (577 of them reasoning) · 8.0 s · $0.0088 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  • DeepSeek V4.1 Flash

    Passed: Graded 4.7 of 5 on average (lowest 4).

    - Aldmoor's six-month Market Street bike lane trial ended; council votes in November.
    - Cyclist trips rose 900→2,300 daily, injuries fell 9→2; car traffic fell 14%, rush-hour delays +1 minute.
    - Businesses split; €410k installed, €1.1m more to make permanent.

    562 tokens in, 310 out (238 of them reasoning) · 2.5 s · $0.0003 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via StreamLake ·

How we ranked them

Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand.

Questions

What is the best AI for research?

In our own test runs on September 2026, on the research prompts: DeepSeek V4.1 Flash, Grok 4.7, Gemini 3.1 Pro, Claude Opus 5.5, and GPT-6 Astra each passed 15 of 15, the most; DeepSeek V4.1 Flash is first on the tie-break (the smaller message, then the lower cost); DeepSeek V4.1 Flash did best of the everyday models (15 of 15), at 60 a day on Pro.

Did the test use web search?

No: every prompt gives the model what to work from (a handbook, a thread, a table), so the test is reading and reasoning, not finding. In llmwise the model can also search the web and cite its sources, which the runs don't use.

Can llmwise write a cited research report?

Research reports, many searches rolled into one cited report, need a paid plan. Web search itself works inside your 5 free messages on the free trial.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.