Skip to content

Best AI · Assistants

Best AI assistants, tested (September 2026)

In our test runs of all 15 models, DeepSeek V4.1 Flash, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 49 of 50, the most; DeepSeek V4.1 Flash is first on more of the hard ones (20 of 20). Below, our ranking across 10 jobs, then 9 AI assistant apps' prices, limits and free tiers in their own words, checked September 28, 2026.

The apps' facts checked against ChatGPT pricing, OpenAI Help Center: What is ChatGPT Plus?, Claude pricing, Gemini subscriptions, Microsoft Support: Microsoft Copilot (free) and Copilot in Microsoft 365, Copilot pricing for individuals, Perplexity pricing, xAI docs: Grok website and apps FAQ, Grok plans, Meta AI assistant FAQ, Meta Newsroom: Introducing Meta One, App Store: DeepSeek - AI Assistant, Vibe pricing. Updated .

Short answer

DeepSeek V4.1 Flash, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 49 of 50, the most; DeepSeek V4.1 Flash is first on more of the hard ones (20 of 20). This is our own test, run on September 27, 2026: every prompt, reply and score is published.

Ranked by our test runs

This is our own test: the same prompts sent to every model through llmwise, each reply checked the same way, with every prompt, reply and score published. It ranks the models, not the apps: we didn't test ChatGPT, Claude, Gemini, Microsoft Copilot, Perplexity, Grok, Meta AI, DeepSeek, and Le Chat (now Vibe) themselves.

Models ranked on coding, writing, math, summarization, data analysis, customer support, translation, SQL, RAG and answering from documents, and agents and tool use
#ModelPassedHard onesFree trialOn ProCost per reply
1DeepSeek V4.1 FlashDeepSeek49 of 5020 of 20Yes60 a day$0.0003
2Gemini 3.1 Pro (preview)Google49 of 5019 of 20Yes125 a month$0.0107
3Claude Opus 5.5Anthropic49 of 5019 of 20No62 a month$0.0107
4GPT-6 AstraOpenAI48 of 5020 of 20No31 a month$0.0117
5GLM 5.3Z.ai48 of 5019 of 20Yes250 a month$0.0007
6GPT-6 LunaOpenAI47 of 5020 of 20Yes60 a day$0.0001
7Grok 4.7xAI47 of 5020 of 20Yes250 a month$0.0063
8Claude Sonnet 5Anthropic46 of 5019 of 20Yes125 a month$0.0041
9Gemini 3.8 FlashGoogle46 of 5018 of 20Yes250 a month$0.0013
10Kimi K3Moonshot46 of 5018 of 20Yes125 a month$0.0039
11GPT-6 SolOpenAI45 of 5019 of 20Yes125 a month$0.0026
12DeepSeek V4 ProDeepSeek45 of 5017 of 20Yes250 a month$0.0021
13Claude Fable 5.1Anthropic45 of 5017 of 20No31 a month$0.0200
14Claude Haiku 4.5Anthropic43 of 5015 of 20Yes250 a month$0.0015
15GLM 5.3 FlashZ.ai40 of 5016 of 20Yes60 a day$0.0002
50 prompts per model (coding, writing, math, summarization, data analysis, customer support, translation, SQL, RAG and answering from documents, and agents and tool use), run on September 27, 2026. Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand. Cost per reply is what OpenRouter charged us on average; in llmwise you pay per message.

The prompts, every reply and how each was scored: our test runs.

The best pick at each price

The same results, by what a message counts as on Pro: the best model at each size of message, then the rest of that size.

  • 60 a day on Pro (everyday models)

    DeepSeek V4.1 Flash, 49 of 50 passed; then GPT-6 Luna 47 of 50 and GLM 5.3 Flash 40 of 50.

  • 125 a month on Pro

    Gemini 3.1 Pro, 49 of 50 passed; then Claude Sonnet 5 46 of 50, Kimi K3 46 of 50, and GPT-6 Sol 45 of 50.

  • 62 a month on Pro

    Claude Opus 5.5, 49 of 50 passed.

  • 31 a month on Pro

    GPT-6 Astra, 48 of 50 passed; then Claude Fable 5.1 45 of 50.

  • 250 a month on Pro

    GLM 5.3, 48 of 50 passed; then Grok 4.7 47 of 50, Gemini 3.8 Flash 46 of 50, DeepSeek V4 Pro 45 of 50, and Claude Haiku 4.5 43 of 50.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Where the top models differ

The top 5, job by job.

Prompts passed on each job, by model
JobDeepSeek V4.1 FlashGemini 3.1 ProClaude Opus 5.5GPT-6 AstraGLM 5.3
Coding5 of 55 of 55 of 55 of 55 of 5
Writing4 of 54 of 54 of 55 of 54 of 5
Math5 of 55 of 55 of 55 of 55 of 5
Summarization5 of 55 of 55 of 55 of 54 of 5
Data analysis5 of 55 of 55 of 55 of 55 of 5
Customer support5 of 55 of 55 of 53 of 55 of 5
Translation5 of 55 of 55 of 55 of 55 of 5
SQL5 of 55 of 55 of 55 of 55 of 5
RAG and answering from documents5 of 55 of 55 of 55 of 55 of 5
Agents and tool use5 of 55 of 55 of 55 of 55 of 5

One prompt, side by side

A hard writing prompt, “Argue both sides of free buses”, where the models didn't all pass.

The prompt: Argue both sides of free buses (hard)
In two paragraphs of at most 90 words each, first argue that a city should make its buses free, then make the strongest case against it. The second paragraph must be as persuasive as the first. No headings.
  • Claude Opus 5.5

    Failed: Graded 4.7 of 5 on average (lowest 4); but a paragraph of 93 words, over the 90 allowed.

    Free buses turn public transit into what it should be: a public good, like sidewalks or streetlights. Fare collection often recovers only a small share of operating costs, yet it slows boarding, requires costly equipment, and deters the riders who need buses most. Remove fares and ridership climbs, as Kansas City and Tallinn saw, easing congestion and cutting emissions. Low-income workers gain hundreds of dollars a year, drivers stop policing payments, and buses move faster. A city that funds roads without tolls can surely fund buses without fares.
    
    Free is not the same as good. Fares fund service, and losing them forces a choice between tax hikes and cuts to frequency, reliability, and coverage, which are the things riders say they want. Studies of fare-free systems find many new riders were pedestrians and cyclists, not drivers, so congestion and emissions barely budge. Some cities report more disruptive incidents, driving away commuters. Low-income riders are better served by targeted discounts, while the revenue preserved buys more buses, more routes, and shorter waits. A bus that never comes is worthless at any price.

    520 tokens in, 2,109 out (1,746 of them reasoning) · 22.5 s · $0.0443 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  • GPT-6 Astra

    Passed: Graded 4.7 of 5 on average (lowest 4).

    A city should make its buses free because mobility is a gateway to work, education, and healthcare—not a privilege reserved for those who can afford each trip. Eliminating fares delivers immediate relief to low-income households, speeds boarding, and removes costly fare enforcement. Funded through progressive taxation, free buses spread the cost fairly while making public transport an easier choice. A city already pays for streets everyone can use; it should support shared travel that uses those streets more efficiently.
    
    But a free bus that arrives late—or never comes—is no bargain. Cities should spend scarce transit funding on frequent, reliable service before subsidizing every rider, including those who can comfortably pay. Lost fare revenue must come from higher taxes, other public services, or the transit budget itself. Targeted discounts can protect people who need help without sacrificing coverage and capacity. To persuade drivers to switch, buses must compete on convenience and reliability, not merely price; better service may deliver greater social and environmental benefits.

    384 tokens in, 197 out · 7.2 s · $0.0137 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

  • DeepSeek V4.1 Flash

    Passed: Graded 4.3 of 5 on average (lowest 4).

    Free buses work. They cut car trips, congestion, and carbon, and they help the riders who need them most. Boarding speeds up, so routes run faster and cheaper to operate. The cost is modest—often less than a parking garage or a lane widening—and downtown shops gain customers. A city that charges nothing at the farebox treats mobility as a right, not a fee.
    
    But somebody still pays. Fares cover a real slice of operating costs, and replacing them means higher taxes or cuts elsewhere—often taxes that hit the poor hardest. Most riders already pay, so the subsidy flows to everyone, including the comfortable. Crowding slows buses, driving away choice riders. Better to means-test discounts, keep service frequent, and spend scarce dollars where they move the most people.

    423 tokens in, 282 out (122 of them reasoning) · 3.9 s · $0.0003 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via StreamLake ·

The AI assistant apps, in their own words

Each app's free tier, cheapest paid plan for one person, and how it words its limits, quoted from its own pages. We didn't test these apps.

AI assistant apps: price, limits and what's free
AppWhat's freeCheapest paid planIts limits, in its words
ChatGPTOpenAI; its GPT models are in llmwise

ChatGPT pricing, OpenAI Help Center: What is ChatGPT Plus?, checked

“The free version of ChatGPT is available to everyone.”ChatGPT Plus“ChatGPT Plus is a subscription plan that provides enhanced access to the ChatGPT web app for $20/month.”“To ensure a smooth experience for all users, Plus subscriptions may include usage limits such as message caps, especially during high demand.”
ClaudeAnthropic; its Claude models are in llmwise

Claude pricing, checked

“Free for everyone”Claude Pro“$20 if billed monthly.”“Pro gives you at least 5x more usage per 5-hour session than Free.”
GeminiGoogle; its Gemini models are in llmwise

Gemini subscriptions, checked

“Varying access to 3.1 Pro”Google AI Pro“Google AI Pro: $19.99/month”“Get 4x higher usage access than Free”
Microsoft CopilotMicrosoft; its models aren't in llmwise

Microsoft Support: Microsoft Copilot (free) and Copilot in Microsoft 365, Copilot pricing for individuals, checked

“Microsoft Copilot is available at no cost at copilot.cloud.microsoft.”Microsoft 365 Personal“Microsoft 365 Personal: $9.99/month”“Overall usage amounts: Higher than free”
PerplexityPerplexity; its models aren't in llmwise

Perplexity pricing, checked

“Good for limited daily usage”Perplexity Pro“$20/month”“Pro raises every limit”
GrokxAI; its Grok models are in llmwise

xAI docs: Grok website and apps FAQ, Grok plans, checked

“You will still have access to Grok's free tier limits on Chat and Voice (these are separate from your weekly usage limit and resets on their own schedule).”SuperGrok“SuperGrok: $30 USD/month”“When you pay for a Grok plan, you receive a weekly usage allowance included with your subscription.”
Meta AIMeta; its models aren't in llmwise

Meta AI assistant FAQ, Meta Newsroom: Introducing Meta One, checked

“Meta AI is free to use for everyday use.”Meta One Core“Core ($7.99/mo)”“generate more images and videos with Meta AI”
DeepSeekDeepSeek; its DeepSeek models are in llmwise

App Store: DeepSeek - AI Assistant, checked

“Experience seamless interaction with DeepSeek's official AI assistant for free!”None listedNot stated
Le Chat (now Vibe)Mistral AI; its models aren't in llmwise

Vibe pricing, checked

“Your personal AI agent for everyday tasks.”Pro“$14.99 /mo”“Messages: Up to 6x free”
llmwiseThis site: 12 of its models are in the trial5 messages on sign-up, once, no card, on 12 models; 1 message with no account when the free-message box shows (GPT-6 Luna answers)Pro, $20 a monthA published count on every model: up to 60 messages a day on the everyday models, and one monthly allowance shown per model

Sources: ChatGPT pricing, OpenAI Help Center: What is ChatGPT Plus?, Claude pricing, Gemini subscriptions, Microsoft Support: Microsoft Copilot (free) and Copilot in Microsoft 365, Copilot pricing for individuals, Perplexity pricing, xAI docs: Grok website and apps FAQ, Grok plans, Meta AI assistant FAQ, Meta Newsroom: Introducing Meta One, App Store: DeepSeek - AI Assistant and Vibe pricing, checked .

How we ranked them

Ranked by how many prompts each model passed, then how many of the hard ones, then by the smaller message (the everyday models first), then by the lower cost per reply. No ranking is chosen by hand.

Questions

What is the best AI assistant?

Among the models, in our own test runs on September 2026: DeepSeek V4.1 Flash, Gemini 3.1 Pro, and Claude Opus 5.5 each passed 49 of 50, the most; DeepSeek V4.1 Flash is first on more of the hard ones (20 of 20), and DeepSeek V4.1 Flash did best of the everyday models (49 of 50). Among the apps, we didn't test them: the table gives each one's price, limits and free tier in its own words, so you can weigh them.

How did you test the AI assistants?

We sent the same 50 prompts, 10 jobs from coding to customer support, to every model llmwise offers, and checked every reply the same way: code by running it, math by its final answer, writing by a grader with a fixed rubric. Every prompt, reply and score is on the methods page.

Did you test ChatGPT, Claude and the other apps?

No. We ran the models llmwise offers, some of which these apps offer too; we didn't use or score the apps themselves, which add their own models, tools, limits and settings. What the table says about each app is quoted from its own page.

Which AI assistant is free?

llmwise's free part is 5 free messages on sign-up, once, with no card; when the free-message box shows, one message needs no account (GPT-6 Luna answers it). Pro is $20 a month.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.