Best AI · Research
The best AI for research, from our test runs
On our research prompts (answering from documents, summarizing and analyzing data), DeepSeek V4.1 Flash, Grok 4.7, Gemini 3.1 Pro, Claude Opus 5.5, and GPT-6 Astra each passed 15 of 15, the most; DeepSeek V4.1 Flash is first on the tie-break (the smaller message, then the lower cost). Every model's result, what each costs, and what the test doesn't cover.
Test runs and prices checked . Updated .
Short answer
DeepSeek V4.1 Flash, Grok 4.7, Gemini 3.1 Pro, Claude Opus 5.5, and GPT-6 Astra each passed 15 of 15, the most; DeepSeek V4.1 Flash is first on the tie-break (the smaller message, then the lower cost). This is our own test, run on September 27, 2026: every prompt, reply and score is published.
Ranked by our test runs
This is our own test: the same prompts sent to every model through llmwise, each reply checked the same way, with every prompt, reply and score published. The research prompts test RAG and answering from documents, summarization, and data analysis over material we give the model; they don't test web search.
| # | Model | Passed | Hard ones | Free trial | On Pro | Cost per reply |
|---|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 FlashDeepSeek | 15 of 15 | 6 of 6 | Yes | 60 a day | $0.0003 |
| 2 | Grok 4.7xAI | 15 of 15 | 6 of 6 | Yes | 250 a month | $0.0047 |
| 3 | Gemini 3.1 Pro (preview)Google | 15 of 15 | 6 of 6 | Yes | 125 a month | $0.0100 |
| 4 | Claude Opus 5.5Anthropic | 15 of 15 | 6 of 6 | No | 62 a month | $0.0100 |
| 5 | GPT-6 AstraOpenAI | 15 of 15 | 6 of 6 | No | 31 a month | $0.0114 |
| 6 | GPT-6 LunaOpenAI | 14 of 15 | 6 of 6 | Yes | 60 a day | $0.0001 |
| 7 | DeepSeek V4 ProDeepSeek | 14 of 15 | 6 of 6 | Yes | 250 a month | $0.0010 |
| 8 | GLM 5.3Z.ai | 14 of 15 | 6 of 6 | Yes | 250 a month | $0.0010 |
| 9 | GPT-6 SolOpenAI | 14 of 15 | 6 of 6 | Yes | 125 a month | $0.0024 |
| 10 | Gemini 3.8 FlashGoogle | 14 of 15 | 5 of 6 | Yes | 250 a month | $0.0014 |
| 11 | Kimi K3Moonshot | 13 of 15 | 6 of 6 | Yes | 125 a month | $0.0039 |
| 12 | Claude Fable 5.1Anthropic | 13 of 15 | 6 of 6 | No | 31 a month | $0.0196 |
| 13 | Claude Haiku 4.5Anthropic | 13 of 15 | 5 of 6 | Yes | 250 a month | $0.0015 |
| 14 | GLM 5.3 FlashZ.ai | 12 of 15 | 5 of 6 | Yes | 60 a day | $0.0002 |
| 15 | Claude Sonnet 5Anthropic | 12 of 15 | 5 of 6 | Yes | 125 a month | $0.0040 |
The prompts, every reply and how each was scored: our test runs.
The best pick at each price
The same results, by what a message counts as on Pro: the best model at each size of message, then the rest of that size.
60 a day on Pro (everyday models)
DeepSeek V4.1 Flash, 15 of 15 passed; then GPT-6 Luna 14 of 15 and GLM 5.3 Flash 12 of 15.
250 a month on Pro
Grok 4.7, 15 of 15 passed; then DeepSeek V4 Pro 14 of 15, GLM 5.3 14 of 15, Gemini 3.8 Flash 14 of 15, and Claude Haiku 4.5 13 of 15.
125 a month on Pro
Gemini 3.1 Pro, 15 of 15 passed; then GPT-6 Sol 14 of 15, Kimi K3 13 of 15, and Claude Sonnet 5 12 of 15.
62 a month on Pro
Claude Opus 5.5, 15 of 15 passed.
31 a month on Pro
GPT-6 Astra, 15 of 15 passed; then Claude Fable 5.1 13 of 15.
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
Where the top models differ
The top 5, job by job.
| Job | DeepSeek V4.1 Flash | Grok 4.7 | Gemini 3.1 Pro | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|---|---|
| RAG and answering from documents | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Summarization | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Data analysis | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
One prompt, side by side
A everyday summarization prompt, “An article in three bullets”, where the models didn't all pass.
The prompt: An article in three bullets (everyday)
Summarize this article in exactly 3 bullet points, at most 60 words in total. The city of Aldmoor ended its six-month trial of protected bike lanes on Market Street last week, and the council will vote in November on whether to keep them. Bike trips on the street rose from about 900 to 2,300 a day during the trial, according to automatic counters. Car traffic fell by 14%, but travel times for drivers rose by just over a minute at rush hour. Shop owners are split: a survey of 60 businesses found 27 reported more customers, 19 reported fewer, and the rest saw no change. Two cyclist injuries were recorded on the street during the trial, down from nine in the same months last year. The lanes cost €410,000 to install; making them permanent would cost another €1.1 million for kerbs and new traffic lights.
Claude Fable 5.1
Failed: Graded 4.3 of 5 on average (lowest 4); but 62 words, over the 60 allowed.
- Aldmoor's six-month protected bike lane trial on Market Street ended; council votes in November on making it permanent. - Daily bike trips rose from 900 to 2,300, car traffic fell 14%, rush-hour driver times rose about a minute, and cyclist injuries dropped from nine to two. - Businesses were split (27 gained customers, 19 lost); permanence would cost €1.1 million beyond the €410,000 installation.
727 tokens in, 150 out · 5.1 s · $0.0164 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·
Gemini 3.1 Pro
Passed: Graded 4.3 of 5 on average (lowest 4).
* Daily bike trips jumped to 2,300 and cyclist injuries dropped sharply. * Car traffic fell 14% with minor delays, while local business impacts were mixed. * The city council votes in November on spending €1.1 million to make the lanes permanent.
560 tokens in, 636 out (577 of them reasoning) · 8.0 s · $0.0088 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·
DeepSeek V4.1 Flash
Passed: Graded 4.7 of 5 on average (lowest 4).
- Aldmoor's six-month Market Street bike lane trial ended; council votes in November. - Cyclist trips rose 900→2,300 daily, injuries fell 9→2; car traffic fell 14%, rush-hour delays +1 minute. - Businesses split; €410k installed, €1.1m more to make permanent.
562 tokens in, 310 out (238 of them reasoning) · 2.5 s · $0.0003 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via StreamLake ·
How we ranked them
Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand.
Questions
What is the best AI for research?
In our own test runs on September 2026, on the research prompts: DeepSeek V4.1 Flash, Grok 4.7, Gemini 3.1 Pro, Claude Opus 5.5, and GPT-6 Astra each passed 15 of 15, the most; DeepSeek V4.1 Flash is first on the tie-break (the smaller message, then the lower cost); DeepSeek V4.1 Flash did best of the everyday models (15 of 15), at 60 a day on Pro.
Did the test use web search?
No: every prompt gives the model what to work from (a handbook, a thread, a table), so the test is reading and reasoning, not finding. In llmwise the model can also search the web and cite its sources, which the runs don't use.
Can llmwise write a cited research report?
Research reports, many searches rolled into one cited report, need a paid plan. Web search itself works inside your 5 free messages on the free trial.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.