Model vs model
DeepSeek V4 Pro vs Grok 4.7
DeepSeek V4 Pro and Grok 4.7 are both in llmwise. What each gets on every plan, what it reads, how it's served and what happens when its provider fails, from the catalog and the code that runs them.
Model prices and specs checked against OpenRouter's DeepSeek V4 Pro page, OpenRouter's Grok 4.7 page. Updated .
Short answer
They cost the same in llmwise: up to 250 messages a month on Pro on either. Otherwise, only Grok 4.7 reads images and only Grok 4.7 reads a PDF as the whole file. In our test runs, DeepSeek V4 Pro passed 45 of the 50 prompts both answered and Grok 4.7 47; 6 prompts split them, most on coding (4 to 5).
DeepSeek V4 Pro vs Grok 4.7, prompt by prompt
Every prompt DeepSeek V4 Pro and Grok 4.7 both answered, compared directly, their biggest differences first. One run each, through OpenRouter: a wait depends on the provider and the load that day, so a lead under 10% counts as close.
Of the 50 prompts both answered, both passed 43, only DeepSeek V4 Pro passed 2, only Grok 4.7 passed 4, and neither passed 1. DeepSeek V4 Pro answered sooner on 36 of the 50 and Grok 4.7 on 11; the rest were within 10% of each other. The 50 replies cost $0.1030 on DeepSeek V4 Pro and $0.3163 on Grok 4.7: 3.1× less on DeepSeek V4 Pro.
The 6 prompts only one of DeepSeek V4 Pro and Grok 4.7 passed
Parse CSV with quoted fields (coding): Grok 4.7 passed and DeepSeek V4 Pro didn't. DeepSeek V4 Pro: No answer within Pro's reply limit of 8,000 tokens: the model spent them all reasoning. Grok 4.7: All 8 tests passed.
Announce a second bakery shop on LinkedIn (writing): DeepSeek V4 Pro passed and Grok 4.7 didn't. DeepSeek V4 Pro: Graded 4.3 of 5 on average (lowest 4). Grok 4.7: Graded 4.5 of 5 on average (lowest 4); but 94 words, under the 100 asked for.
Rewrite corporate jargon in plain words (writing): DeepSeek V4 Pro passed and Grok 4.7 didn't. DeepSeek V4 Pro: Graded 4.0 of 5 on average (lowest 3). Grok 4.7: Graded 3.7 of 5 on average (lowest 3).
A product announcement with five rules (writing): Grok 4.7 passed and DeepSeek V4 Pro didn't. DeepSeek V4 Pro: Graded 3.7 of 5 on average (lowest 2); but doesn't end with a question. Grok 4.7: Graded 4.0 of 5 on average (lowest 3).
Argue both sides of free buses (writing): Grok 4.7 passed and DeepSeek V4 Pro didn't. DeepSeek V4 Pro: Graded 4.0 of 5 on average (lowest 4); but a paragraph of 95 words, over the 90 allowed. Grok 4.7: Graded 4.3 of 5 on average (lowest 4).
An email thread in one sentence (summarization): Grok 4.7 passed and DeepSeek V4 Pro didn't. DeepSeek V4 Pro: Graded 4.7 of 5 on average (lowest 4); but 33 words, over the 30 allowed. Grok 4.7: Graded 4.3 of 5 on average (lowest 4).
Job by job, the widest gaps first
Summarization: DeepSeek V4 Pro passed 4 of 5 and Grok 4.7 5 of 5. DeepSeek V4 Pro answered 1.5× sooner at the median, 3.1 s against 4.7 s. DeepSeek V4 Pro cost 3.6× less, $0.0045 against $0.0161 for the 5 replies. Their replies ran to about the same length.
Coding: DeepSeek V4 Pro passed 4 of 5 and Grok 4.7 5 of 5. DeepSeek V4 Pro answered 1.3× sooner at the median, 20.1 s against 25.3 s. DeepSeek V4 Pro cost 1.7× less, $0.0662 against $0.1099 for the 5 replies. Their replies ran to about the same length.
Agents and tool use: DeepSeek V4 Pro passed 5 of 5 and Grok 4.7 5 of 5. Grok 4.7 answered 1.3× sooner at the median, 3.3 s against 2.5 s. DeepSeek V4 Pro cost 9.5× less, $0.0014 against $0.0129 for the 5 replies. DeepSeek V4 Pro's replies ran 12% longer, in tokens of reply, thinking not counted.
Writing: DeepSeek V4 Pro passed 3 of 5 and Grok 4.7 3 of 5. DeepSeek V4 Pro answered 1.8× sooner at the median, 2.2 s against 4.0 s. DeepSeek V4 Pro cost 9.4× less, $0.0022 against $0.0210 for the 5 replies. DeepSeek V4 Pro's replies ran 34% longer, in tokens of reply, thinking not counted.
Translation: DeepSeek V4 Pro passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4 Pro answered 3.2× sooner at the median, 4.3 s against 13.8 s. DeepSeek V4 Pro cost 7.7× less, $0.0045 against $0.0347 for the 5 replies. DeepSeek V4 Pro's replies ran 20% longer, in tokens of reply, thinking not counted.
Math: DeepSeek V4 Pro passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4 Pro answered 3.3× sooner at the median, 2.7 s against 9.1 s. DeepSeek V4 Pro cost 7.6× less, $0.0033 against $0.0251 for the 5 replies. Grok 4.7's replies ran 34% longer, in tokens of reply, thinking not counted.
RAG and answering from documents: DeepSeek V4 Pro passed 5 of 5 and Grok 4.7 5 of 5. Their median waits were close, 1.9 s against 1.9 s. DeepSeek V4 Pro cost 6.8× less, $0.0017 against $0.0115 for the 5 replies. DeepSeek V4 Pro's replies ran 13% longer, in tokens of reply, thinking not counted.
Data analysis: DeepSeek V4 Pro passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4 Pro answered 3.8× sooner at the median, 2.5 s against 9.4 s. DeepSeek V4 Pro cost 5.2× less, $0.0081 against $0.0423 for the 5 replies. Their replies ran to about the same length.
Customer support: DeepSeek V4 Pro passed 4 of 5 and Grok 4.7 4 of 5. DeepSeek V4 Pro answered 1.6× sooner at the median, 5.1 s against 7.9 s. DeepSeek V4 Pro cost 3.9× less, $0.0061 against $0.0239 for the 5 replies. Their replies ran to about the same length.
SQL: DeepSeek V4 Pro passed 5 of 5 and Grok 4.7 5 of 5. DeepSeek V4 Pro answered 1.3× sooner at the median, 3.8 s against 4.9 s. DeepSeek V4 Pro cost 3.8× less, $0.0049 against $0.0189 for the 5 replies. DeepSeek V4 Pro's replies ran 21% longer, in tokens of reply, thinking not counted.
All 50 prompts: who passed, who answered sooner, who cost less
| Prompt | Result | Sooner | Cheaper |
|---|---|---|---|
| Turn a title into a URL slug | Both passed | DeepSeek V4 Pro, 3.2×took 2.1 s and 6.7 s | DeepSeek V4 Pro, 3.7×cost $0.0013 and $0.0048 |
| Parse a duration like “1h 30m” | Both passed | DeepSeek V4 Pro, 1.3×took 20.1 s and 25.3 s | DeepSeek V4 Pro, 5.1×cost $0.0030 and $0.0151 |
| Merge overlapping intervals | Both passed | Grok 4.7, 1.7×took 6.6 s and 3.8 s | DeepSeek V4 Pro, 5.1×cost $0.0005 and $0.0025 |
| Evaluate an arithmetic expression, no eval | Both passed | DeepSeek V4 Pro, 3.1×took 40.3 s and 124.9 s | DeepSeek V4 Pro, 1.6×cost $0.0291 and $0.0472 |
| Parse CSV with quoted fields | Only Grok 4.7 | DeepSeek V4 Pro, 1.5×took 69.1 s and 100.4 s | DeepSeek V4 Pro, 1.2×cost $0.0323 and $0.0403 |
| Announce a second bakery shop on LinkedIn | Only DeepSeek V4 Pro | DeepSeek V4 Pro, 2.1×took 1.9 s and 4.0 s | DeepSeek V4 Pro, 2.6×cost $0.0008 and $0.0022 |
| Rewrite corporate jargon in plain words | Only DeepSeek V4 Pro | DeepSeek V4 Pro, 16.7×took 1.2 s and 19.3 s | DeepSeek V4 Pro, 14.8×cost $0.0006 and $0.0088 |
| Decline a meeting and offer two times | Both passed | Closetook 2.2 s and 2.3 s | DeepSeek V4 Pro, 5.6×cost $0.0003 and $0.0019 |
| A product announcement with five rules | Only Grok 4.7 | Grok 4.7, 1.5×took 3.8 s and 2.6 s | DeepSeek V4 Pro, 11.5×cost $0.0001 and $0.0017 |
| Argue both sides of free buses | Only Grok 4.7 | DeepSeek V4 Pro, 1.5×took 7.8 s and 12.1 s | DeepSeek V4 Pro, 20.7×cost $0.0003 and $0.0065 |
| A discount, then sales tax | Both passed | DeepSeek V4 Pro, 1.4×took 2.7 s and 3.9 s | DeepSeek V4 Pro, 14.1×cost $0.0002 and $0.0027 |
| Pens at 3 for $4 | Both passed | DeepSeek V4 Pro, 15.7×took 2.6 s and 40.7 s | DeepSeek V4 Pro, 5.6×cost $0.0016 and $0.0087 |
| Compound interest over three years | Both passed | DeepSeek V4 Pro, 3.6×took 1.5 s and 5.5 s | DeepSeek V4 Pro, 8.9×cost $0.0004 and $0.0031 |
| Four-digit numbers whose digits sum to 9 | Both passed | DeepSeek V4 Pro, 3.5×took 3.2 s and 11.1 s | DeepSeek V4 Pro, 9.3×cost $0.0006 and $0.0056 |
| The highest of three dice is a 5 | Both passed | DeepSeek V4 Pro, 2.6×took 3.5 s and 9.1 s | DeepSeek V4 Pro, 8.2×cost $0.0006 and $0.0050 |
| An article in three bullets | Both passed | DeepSeek V4 Pro, 1.6×took 6.1 s and 9.8 s | DeepSeek V4 Pro, 2.5×cost $0.0021 and $0.0051 |
| An email thread in one sentence | Only Grok 4.7 | Grok 4.7, 1.6×took 2.8 s and 1.7 s | DeepSeek V4 Pro, 12.3×cost $0.0001 and $0.0018 |
| Decisions and action items from a meeting | Both passed | DeepSeek V4 Pro, 3.2×took 1.5 s and 4.7 s | DeepSeek V4 Pro, 5.8×cost $0.0005 and $0.0028 |
| A quarterly memo for the CEO | Both passed | DeepSeek V4 Pro, 2.0×took 3.1 s and 6.1 s | DeepSeek V4 Pro, 24.5×cost $0.0002 and $0.0044 |
| A study with a negative result | Both passed | Grok 4.7, 1.2×took 3.5 s and 3.0 s | DeepSeek V4 Pro, 1.2×cost $0.0016 and $0.0020 |
| The region with the most revenue | Both passed | DeepSeek V4 Pro, 4.8×took 2.0 s and 9.4 s | DeepSeek V4 Pro, 4.0×cost $0.0014 and $0.0057 |
| Average order value in August | Both passed | DeepSeek V4 Pro, 1.5×took 3.7 s and 5.5 s | DeepSeek V4 Pro, 1.8×cost $0.0021 and $0.0037 |
| Revenue change from July to August | Both passed | DeepSeek V4 Pro, 4.1×took 2.5 s and 10.1 s | DeepSeek V4 Pro, 4.2×cost $0.0017 and $0.0071 |
| A median, filtered two ways | Both passed | DeepSeek V4 Pro, 3.2×took 1.3 s and 4.2 s | DeepSeek V4 Pro, 4.2×cost $0.0008 and $0.0036 |
| Correlation between ad spend and sign-ups | Both passed | DeepSeek V4 Pro, 1.1×took 41.2 s and 45.4 s | DeepSeek V4 Pro, 10.4×cost $0.0021 and $0.0222 |
| A late order | Both passed | Grok 4.7, 1.2×took 20.2 s and 17.2 s | DeepSeek V4 Pro, 1.8×cost $0.0041 and $0.0076 |
| A return inside the window | Both passed | DeepSeek V4 Pro, 2.0×took 2.4 s and 4.7 s | DeepSeek V4 Pro, 3.4×cost $0.0009 and $0.0032 |
| A frustrated customer | Neither passed | DeepSeek V4 Pro, 1.1×took 8.6 s and 9.4 s | DeepSeek V4 Pro, 9.9×cost $0.0005 and $0.0050 |
| A refund request outside the window | Both passed | DeepSeek V4 Pro, 2.0×took 3.9 s and 7.6 s | DeepSeek V4 Pro, 15.1×cost $0.0003 and $0.0041 |
| A message with a planted instruction | Both passed | DeepSeek V4 Pro, 1.6×took 5.1 s and 7.9 s | DeepSeek V4 Pro, 12.7×cost $0.0003 and $0.0040 |
| A delivery message into Spanish | Both passed | DeepSeek V4 Pro, 3.2×took 4.3 s and 13.8 s | DeepSeek V4 Pro, 7.2×cost $0.0009 and $0.0066 |
| A product description into French | Both passed | DeepSeek V4 Pro, 7.3×took 3.1 s and 22.7 s | DeepSeek V4 Pro, 43.4×cost $0.0002 and $0.0097 |
| A meeting note into German | Both passed | Grok 4.7, 1.7×took 22.0 s and 12.6 s | DeepSeek V4 Pro, 6.0×cost $0.0010 and $0.0060 |
| Idioms into natural Japanese | Both passed | DeepSeek V4 Pro, 2.1×took 4.7 s and 10.0 s | DeepSeek V4 Pro, 3.2×cost $0.0015 and $0.0048 |
| A lease clause into Brazilian Portuguese | Both passed | DeepSeek V4 Pro, 4.0×took 3.7 s and 15.0 s | DeepSeek V4 Pro, 9.0×cost $0.0008 and $0.0076 |
| Customers in one country | Both passed | Closetook 1.9 s and 1.8 s | DeepSeek V4 Pro, 10.7×cost $0.0002 and $0.0018 |
| Count orders by status | Both passed | Grok 4.7, 1.9×took 2.7 s and 1.4 s | DeepSeek V4 Pro, 8.8×cost $0.0002 and $0.0017 |
| Revenue by category | Both passed | DeepSeek V4 Pro, 1.3×took 3.8 s and 4.9 s | DeepSeek V4 Pro, 15.3×cost $0.0002 and $0.0031 |
| Every customer, even those without orders | Both passed | DeepSeek V4 Pro, 1.2×took 4.5 s and 5.3 s | DeepSeek V4 Pro, 1.3×cost $0.0028 and $0.0037 |
| Monthly revenue with a running total | Both passed | Grok 4.7, 2.0×took 34.4 s and 16.9 s | DeepSeek V4 Pro, 5.5×cost $0.0016 and $0.0086 |
| A fact from one section | Both passed | DeepSeek V4 Pro, 1.2×took 2.2 s and 2.7 s | DeepSeek V4 Pro, 10.3×cost $0.0002 and $0.0025 |
| Core hours and start times | Both passed | Grok 4.7, 1.3×took 2.5 s and 1.9 s | DeepSeek V4 Pro, 3.9×cost $0.0006 and $0.0022 |
| Two sections in one answer | Both passed | DeepSeek V4 Pro, 1.1×took 1.3 s and 1.5 s | DeepSeek V4 Pro, 6.1×cost $0.0003 and $0.0018 |
| A later amendment changes the answer | Both passed | DeepSeek V4 Pro, 4.1×took 1.3 s and 5.5 s | DeepSeek V4 Pro, 8.3×cost $0.0004 and $0.0030 |
| A question the handbook doesn't answer | Both passed | Closetook 1.9 s and 1.8 s | DeepSeek V4 Pro, 9.1×cost $0.0002 and $0.0021 |
| Pick the tool and work out the date | Both passed | Grok 4.7, 2.0×took 3.2 s and 1.5 s | DeepSeek V4 Pro, 9.2×cost $0.0002 and $0.0018 |
| Convert a currency | Both passed | Grok 4.7, 2.2×took 3.4 s and 1.5 s | DeepSeek V4 Pro, 12.7×cost $0.0001 and $0.0013 |
| Book a meeting from a sentence | Both passed | DeepSeek V4 Pro, 2.6×took 3.5 s and 9.1 s | DeepSeek V4 Pro, 18.8×cost $0.0003 and $0.0048 |
| Search, but don't book | Both passed | DeepSeek V4 Pro, 1.9×took 1.3 s and 2.5 s | DeepSeek V4 Pro, 3.3×cost $0.0006 and $0.0020 |
| Two calls with a unit conversion | Both passed | DeepSeek V4 Pro, 1.3×took 3.3 s and 4.5 s | DeepSeek V4 Pro, 15.6×cost $0.0002 and $0.0029 |
DeepSeek V4 Pro vs Grok 4.7 in our test runs
DeepSeek V4 Pro and Grok 4.7 on the same prompts, job by job: how many replies passed their check.
Based on 100 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
| Job | DeepSeek V4 Pro | Grok 4.7 |
|---|---|---|
| Coding | 4 of 5 | 5 of 5 |
| Writing | 3 of 5 | 3 of 5 |
| Math | 5 of 5 | 5 of 5 |
| Summarization | 4 of 5 | 5 of 5 |
| Data analysis | 5 of 5 | 5 of 5 |
| Customer support | 4 of 5 | 4 of 5 |
| Translation | 5 of 5 | 5 of 5 |
| SQL | 5 of 5 | 5 of 5 |
| RAG and answering from documents | 5 of 5 | 5 of 5 |
| Agents and tool use | 5 of 5 | 5 of 5 |
| All jobs | 45 of 50 · 1 message each · $0.0021 a reply | 47 of 50 · 1 message each · $0.0063 a reply |
The same prompts, sent to both through llmwise's own pipeline and checked the same way: five for each job. A bold count is the one that passed more. Messages: what each reply counted as on Pro. Cost: what OpenRouter charged us per reply, on average. Every prompt, and how it's scored.
Where they differed most, coding: both replies to a prompt only one of them passed.
A hard prompt: “Parse CSV with quoted fields”, in full.
DeepSeek V4 Pro
Failed: No answer within Pro's reply limit of 8,000 tokens: the model spent them all reasoning.
(no reply)
493 tokens in, 8,000 out (7,805 of them reasoning) · 69.1 s · $0.0323 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via CoreWeave ·
Grok 4.7
Passed: All 8 tests passed.
```javascript function parseCsv(text) { if (text === "") return []; const rows = []; let row = []; let field = ""; let inQuotes = false; let i = 0; while (i < text.length) { const c = text[i]; if (inQuotes) { if (c === '"') { if (text[i + 1] === '"') { field += '"'; i += 2; } else { inQuotes = false; i += 1; } } else { field += c; i += 1; } continue; } if (c === '"') { if (field.length === 0) inQuotes = true; else field += '"'; i += 1; continue; } if (c === ",") { row.push(field); field = ""; i += 1; continue; } if (c === "\n" || (c === "\r" && text[i + 1] === "\n")) { i += c === "\r" ? 2 : 1; row.push(field); rows.push(row); row = []; field = ""; continue; } field += c; i += 1; } if (text[text.length - 1] !== "\n" || inQuotes) { row.push(field); rows.push(row); } return rows; } ```1,717 tokens in, 8,172 out (7,825 of them reasoning) · 100.4 s · $0.0403 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·
DeepSeek V4 Pro and Grok 4.7 on every plan
Whether the one-time free trial reaches each model, then each paid plan's messages on it.
| Plan | Price | DeepSeek V4 Pro | Grok 4.7 |
|---|---|---|---|
| Free | $0 | In the one-time trial of 5 messages | In the one-time trial of 5 messages |
| Pro | $20 a month | Up to 250 a month | Up to 250 a month |
| Max | $50 a month | Up to 800 a month | Up to 800 a month |
| Ultra | $100 a month | Up to 1,800 a month | Up to 1,800 a month |
| Studio | $200 a month | Up to 4,000 a month | Up to 4,000 a month |
Prices don't include tax, which is added where it applies and shown before you pay. A paid plan's month is one allowance shared by every model, so each monthly count is the most you get if all of it goes to that model. It renews each billing period; everyday models refill daily at 00:00 UTC. Long chats count more per reply. How pricing works.
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
What differs
Messages on Pro
They cost the same in llmwise: up to 250 messages a month on Pro on either.
Context window
DeepSeek V4 Pro takes up to 1.05M tokens; Grok 4.7 up to 500K tokens. A chat in llmwise holds up to 200k tokens, which fits in either, so the difference shows only through each maker's own API.
Images and PDFs
DeepSeek V4 Pro doesn't read images. DeepSeek V4 Pro gets a PDF's text rather than the file itself.
On Free
Both are in the free trial.
Where messages go
DeepSeek V4 Pro: Served through OpenRouter, only by hosts that don't store or train on prompts. The maker's own endpoint is never asked. Grok 4.7: Served through OpenRouter by xAI alone, on an endpoint that doesn't store or train on prompts.
Fact by fact
| Fact | DeepSeek V4 Pro | Grok 4.7 |
|---|---|---|
| Context window | 1.05M tokens | 500K tokens |
| Reads images | No | Yes |
| PDFs | Text only | Whole file |
| Reasoning | Yes | Yes |
| API price (September 2026) | $0.44 in / $2.90 out per million tokens | $1.60 in / $4.80 out per million tokens |
| A typical message at API prices (4,000 tokens in, 700 out) | $0.0038 | $0.0098 |
| A $10 top-up adds | 200 messages | 200 messages |
| Where a message goes | Served through OpenRouter, only by hosts that don't store or train on prompts. The maker's own endpoint is never asked. | Served through OpenRouter by xAI alone, on an endpoint that doesn't store or train on prompts. |
| If the provider fails | When one host is down, OpenRouter moves the request to another host that meets the same rules. | xAI is its only host, so there's no other host to move to: if xAI fails, send the message again or pick another model. |
| Anthropic's safety fallback | Doesn't apply | Doesn't apply |
DeepSeek V4 Pro or Grok 4.7?
From the facts above and our test runs: the rest is how their answers suit your work, which one chat can show you.
Pick DeepSeek V4 Pro: it costs its maker less to run ($0.0038 a typical message at API prices), though in llmwise the count is the same.
Pick Grok 4.7: it reads images; it reads a PDF as the whole file, charts and scans included.
Each model's page, the families, and other pairs
DeepSeek V4 Pro vs Grok 4.7 is one pair of models. The page below covers the whole families.
- DeepSeek V4 Pro: price, limits and messages on every plan
- Grok 4.7: price, limits and messages on every plan
- Grok vs DeepSeek
- Claude Opus 5.5 vs DeepSeek V4 Pro
- DeepSeek V4 Pro vs Kimi K3
- DeepSeek V4 Pro vs GLM 5.3
- Claude Fable 5.1 vs Grok 4.7
- GPT-6 Astra vs Grok 4.7
- Claude Sonnet 5 vs Grok 4.7
- Every model-vs-model page
DeepSeek V4 Pro is a DeepSeek model; Grok 4.7 is a Grok model.
Questions
Is DeepSeek V4 Pro or Grok 4.7 cheaper in llmwise?
They cost the same in llmwise: up to 250 messages a month on Pro on either. Every paid plan's monthly allowance is shared by all models, so each count is the most you get if it all goes to that model.
Can I try DeepSeek V4 Pro and Grok 4.7 for free?
Yes: both are in the free trial of 5 messages.
Which has the bigger context window, DeepSeek V4 Pro or Grok 4.7?
DeepSeek V4 Pro: 1.05M tokens, against 500K tokens. A chat in llmwise holds up to 200k tokens, which fits in either, so the difference shows only through each maker's own API.
Can I use DeepSeek V4 Pro and Grok 4.7 in the same chat?
Yes. Pick DeepSeek V4 Pro for one message and Grok 4.7 for the next; the second sees the whole chat, including the first one's answer.
Which did better in your test runs, DeepSeek V4 Pro or Grok 4.7?
On the same 50 prompts, run on September 27, 2026, DeepSeek V4 Pro passed 45 and Grok 4.7 passed 47. The table on this page has each job, and every reply is published.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.