Skip to content

AI data

Updated · By , AI-assisted.

AI model quality tracker: the same prompts, every model, dated

A dated snapshot, not a trend yet: all 19 models ran the same 50 prompts on September 27, 28, 29 and October 2, 7, 8, 9, 2026. They passed from 40 (GLM 5.3 Flash) to 49 (DeepSeek V4 Pro), and all passed the same 33. We rerun within 45 days, or when a price or version changes, keeping each run.

Every model, dated

Each model's run: the day, the version that answered (as OpenRouter reported it), what it passed of our 50 prompts, and the day that run stops counting. Ranked by our published rule.

Every model's run of our test prompts, dated
ModelRan onVersion that answeredPassedHard onesCounts until
DeepSeek V4 ProOctober 8, 2026deepseek/deepseek-v4-pro-081349 of 5020 of 20November 22, 2026
Gemini 3.1 ProSeptember 27, 2026google/gemini-3.1-pro-preview49 of 5019 of 20November 11, 2026
Claude Opus 5.5September 27, 2026anthropic/claude-opus-5.549 of 5019 of 20November 11, 2026
GPT-6 AstraSeptember 27, 2026openai/gpt-6-astra48 of 5020 of 20November 11, 2026
DeepSeek V4.1 FlashOctober 8, 2026deepseek/deepseek-v4.1-flash48 of 5019 of 20November 22, 2026
GLM 5.3September 27, 2026z-ai/glm-5.348 of 5019 of 20November 11, 2026
GPT-6 LunaSeptember 27, 2026openai/gpt-6-luna47 of 5020 of 20November 11, 2026
Grok 4.7October 2, 2026x-ai/grok-4.747 of 5019 of 20November 16, 2026
Claude Sonnet 5.5September 28, 2026anthropic/claude-sonnet-5.547 of 5019 of 20November 12, 2026
Mistral Large 4October 9, 2026mistralai/mistral-large-4-047 of 5017 of 20November 23, 2026
GPT-6.1 SolSeptember 29, 2026openai/gpt-6.1-sol46 of 5019 of 20November 13, 2026
Claude Sonnet 5September 27, 2026anthropic/claude-sonnet-546 of 5019 of 20November 11, 2026
Gemini 3.8 FlashSeptember 27, 2026google/gemini-3.8-flash46 of 5018 of 20November 11, 2026
Kimi K3September 27, 2026moonshotai/kimi-k346 of 5018 of 20November 11, 2026
Claude Haiku 5.5October 7, 2026anthropic/claude-haiku-5.545 of 5019 of 20November 21, 2026
GPT-6 SolSeptember 27, 2026openai/gpt-6-sol45 of 5019 of 20November 11, 2026
Claude Fable 5.1September 27, 2026anthropic/claude-fable-5.145 of 5017 of 20November 11, 2026
Claude Haiku 4.5September 27, 2026anthropic/claude-haiku-4.543 of 5015 of 20November 11, 2026
GLM 5.3 FlashSeptember 27, 2026z-ai/glm-5.3-flash40 of 5016 of 20November 11, 2026

Every prompt, every model

Each prompt with how many models passed it and which didn't: when a model changes, these are the rows that move first.

Each test prompt: how many models passed, and which failed
PromptJobPassedFailed by
Turn a title into a URL slugCoding, everyday19 of 19none
Parse a duration like “1h 30m”Coding, everyday19 of 19none
Merge overlapping intervalsCoding, everyday19 of 19none
Evaluate an arithmetic expression, no evalCoding, hard17 of 19Claude Haiku 4.5, Mistral Large 4
Parse CSV with quoted fieldsCoding, hard18 of 19Claude Haiku 4.5
Announce a second bakery shop on LinkedInWriting, everyday18 of 19Claude Sonnet 5
Rewrite corporate jargon in plain wordsWriting, everyday12 of 19Claude Haiku 5.5, GPT-6.1 Sol, GPT-6 Sol, Gemini 3.8 Flash, DeepSeek V4.1 Flash, Grok 4.7, GLM 5.3 Flash
Decline a meeting and offer two timesWriting, everyday19 of 19none
A product announcement with five rulesWriting, hard13 of 19Claude Fable 5.1, Claude Haiku 4.5, Gemini 3.8 Flash, Kimi K3, GLM 5.3 Flash, Mistral Large 4
Argue both sides of free busesWriting, hard11 of 19Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, Claude Haiku 5.5, Gemini 3.1 Pro, Grok 4.7, GLM 5.3, GLM 5.3 Flash
A discount, then sales taxMath, everyday19 of 19none
Pens at 3 for $4Math, everyday18 of 19GPT-6 Luna
Compound interest over three yearsMath, everyday19 of 19none
Four-digit numbers whose digits sum to 9Math, hard19 of 19none
The highest of three dice is a 5Math, hard19 of 19none
An article in three bulletsSummarization, everyday11 of 19Claude Fable 5.1, Claude Sonnet 5, Claude Haiku 5.5, GPT-6 Sol, GPT-6 Luna, Kimi K3, GLM 5.3, GLM 5.3 Flash
An email thread in one sentenceSummarization, everyday13 of 19Claude Fable 5.1, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Kimi K3, GLM 5.3 Flash
Decisions and action items from a meetingSummarization, everyday19 of 19none
A quarterly memo for the CEOSummarization, hard19 of 19none
A study with a negative resultSummarization, hard19 of 19none
The region with the most revenueData analysis, everyday19 of 19none
Average order value in AugustData analysis, everyday18 of 19Claude Haiku 4.5
Revenue change from July to AugustData analysis, everyday19 of 19none
A median, filtered two waysData analysis, hard19 of 19none
Correlation between ad spend and sign-upsData analysis, hard15 of 19Claude Sonnet 5, Claude Haiku 4.5, Gemini 3.8 Flash, GLM 5.3 Flash
A late orderCustomer support, everyday15 of 19GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, GLM 5.3 Flash
A return inside the windowCustomer support, everyday19 of 19none
A frustrated customerCustomer support, everyday8 of 19Claude Sonnet 5.5, Claude Haiku 5.5, Claude Haiku 4.5, GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, GPT-6 Luna, Gemini 3.8 Flash, DeepSeek V4 Pro, Grok 4.7, GLM 5.3 Flash
A refund request outside the windowCustomer support, hard15 of 19GPT-6.1 Sol, Kimi K3, GLM 5.3 Flash, Mistral Large 4
A message with a planted instructionCustomer support, hard16 of 19Claude Fable 5.1, Claude Haiku 4.5, DeepSeek V4.1 Flash
A delivery message into SpanishTranslation, everyday19 of 19none
A product description into FrenchTranslation, everyday19 of 19none
A meeting note into GermanTranslation, everyday19 of 19none
Idioms into natural JapaneseTranslation, hard18 of 19GPT-6 Sol
A lease clause into Brazilian PortugueseTranslation, hard19 of 19none
Customers in one countrySQL, everyday19 of 19none
Count orders by statusSQL, everyday19 of 19none
Revenue by categorySQL, everyday19 of 19none
Every customer, even those without ordersSQL, hard19 of 19none
Monthly revenue with a running totalSQL, hard19 of 19none
A fact from one sectionRAG and answering from documents, everyday19 of 19none
Core hours and start timesRAG and answering from documents, everyday19 of 19none
Two sections in one answerRAG and answering from documents, everyday19 of 19none
A later amendment changes the answerRAG and answering from documents, hard19 of 19none
A question the handbook doesn't answerRAG and answering from documents, hard19 of 19none
Pick the tool and work out the dateAgents and tool use, everyday19 of 19none
Convert a currencyAgents and tool use, everyday19 of 19none
Book a meeting from a sentenceAgents and tool use, everyday18 of 19GLM 5.3 Flash
Search, but don't bookAgents and tool use, hard19 of 19none
Two calls with a unit conversionAgents and tool use, hard19 of 19none

Every prompt in full, and how each reply is checked.

More AI data

Also from primary sources: OpenRouter's usage accounting docs (the cost OpenRouter reports for every request, which these figures add up). Anthropic's Claude Opus 5.5 page (the model that grades the rubric prompts). OpenAI's GPT-6 Astra docs (the model that grades Claude Opus 5.5's own replies). Read .

Bar chart: Prompts passed in our test runs, all 10 jobs. DeepSeek V4 Pro: 49 of 50; Gemini 3.1 Pro (preview): 49 of 50; Claude Opus 5.5: 49 of 50; GPT-6 Astra: 48 of 50; DeepSeek V4.1 Flash: 48 of 50; GLM 5.3: 48 of 50; GPT-6 Luna: 47 of 50; Grok 4.7: 47 of 50; Claude Sonnet 5.5: 47 of 50; Mistral Large 4: 47 of 50; GPT-6.1 Sol: 46 of 50; Claude Sonnet 5: 46 of 50; Gemini 3.8 Flash: 46 of 50; Kimi K3: 46 of 50; Claude Haiku 5.5: 45 of 50; GPT-6 Sol: 45 of 50; Claude Fable 5.1: 45 of 50; Claude Haiku 4.5: 43 of 50; GLM 5.3 Flash: 40 of 50.
Our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026: the same prompts for every model, each reply checked the same way.

The quality tracker: method

Has a model got worse?

Not from these runs yet: each model has run the prompts once, so there's nothing earlier to hold them against. Each re-run will be kept and shown here beside this one, with the version that answered, so a change shows as a change.

How often are the prompts run again?

Within 45 days of the last run, and sooner when a model's price or version changes in our catalog: a run of a model at another price or version no longer counts, and the pages built on it say so.

What's tested?

The same prompts, sent the way llmwise sends a message, checked the same way for every model: tests that run the code, the final answer for math and data analysis, the query's result for SQL, the key facts for documents, the tool calls for agents, a grader model's published rubric for writing, summaries and support, and back-translation for translation.

Check a model on your own question

The tracker scores 50 fixed prompts. Your questions may rank the models differently: ask two of them the same thing in one llmwise chat.