Skip to content

Comparison

Grok vs GLM

On llmwise Pro, Grok 4.7 up to 250 messages a month and GLM 5.3 up to 250. We ran the same 50 prompts across 10 jobs on every Grok and GLM model and published every reply: each family's pick against the other's job by job, the prompts where they split, then their plans and how the lineups differ.

Based on 150 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: Grok's one model passed 47 of 50 replies and GLM's 2 models passed 88 of 100 replies. Job by job, both families' best models shared the top result on 7 of the 10 jobs; Grok's alone had it on summarization; GLM's alone had it on writing and customer support.

Grok vs GLM, job by job

On each job, Grok's pick against GLM's: the model of each family that passed the most of the job's 5 prompts (then the most hard ones, then the cheaper). The jobs where they differ most come first.

  • Writing: Grok 4.7 passed 3 of 5 and GLM 5.3 4 of 5. GLM 5.3 answered 2.3× sooner at the median, 4.0 s against 1.8 s. GLM 5.3 cost 7.3× less, $0.0210 against $0.0029 for the 5 replies. GLM 5.3's replies ran 41% longer, in tokens of reply, thinking not counted.

  • Summarization: Grok 4.7 passed 5 of 5 and GLM 5.3 4 of 5. GLM 5.3 answered 4.3× sooner at the median, 4.7 s against 1.1 s. GLM 5.3 cost 5.9× less, $0.0161 against $0.0027 for the 5 replies. Their replies ran to about the same length.

  • Customer support: Grok 4.7 passed 4 of 5 and GLM 5.3 5 of 5. GLM 5.3 answered 3.4× sooner at the median, 7.9 s against 2.3 s. GLM 5.3 cost 10.3× less, $0.0239 against $0.0023 for the 5 replies. Their replies ran to about the same length.

  • Coding: Grok 4.7 passed 5 of 5 and GLM 5.3 Flash 5 of 5. GLM 5.3 Flash answered 3.0× sooner at the median, 25.3 s against 8.4 s. GLM 5.3 Flash cost 86.9× less, $0.1099 against $0.0013 for the 5 replies. Their replies ran to about the same length.

  • Math: Grok 4.7 passed 5 of 5 and GLM 5.3 Flash 5 of 5. GLM 5.3 Flash answered 3.3× sooner at the median, 9.1 s against 2.7 s. GLM 5.3 Flash cost 33.5× less, $0.0251 against $0.0008 for the 5 replies. Grok 4.7's replies ran 13% longer, in tokens of reply, thinking not counted.

  • Data analysis: Grok 4.7 passed 5 of 5 and GLM 5.3 5 of 5. GLM 5.3 answered 4.8× sooner at the median, 9.4 s against 1.9 s. GLM 5.3 cost 4.4× less, $0.0423 against $0.0096 for the 5 replies. GLM 5.3's replies ran 24% longer, in tokens of reply, thinking not counted.

  • Translation: Grok 4.7 passed 5 of 5 and GLM 5.3 Flash 5 of 5. GLM 5.3 Flash answered 3.4× sooner at the median, 13.8 s against 4.1 s. GLM 5.3 Flash cost 70.2× less, $0.0347 against $0.0005 for the 5 replies. GLM 5.3 Flash's replies ran 24% longer, in tokens of reply, thinking not counted.

  • SQL: Grok 4.7 passed 5 of 5 and GLM 5.3 Flash 5 of 5. GLM 5.3 Flash answered 3.4× sooner at the median, 4.9 s against 1.4 s. GLM 5.3 Flash cost 29.2× less, $0.0189 against $0.0006 for the 5 replies. Their replies ran to about the same length.

  • RAG and answering from documents: Grok 4.7 passed 5 of 5 and GLM 5.3 Flash 5 of 5. Grok 4.7 answered 1.2× sooner at the median, 1.9 s against 2.2 s. GLM 5.3 Flash cost 15.0× less, $0.0115 against $0.0008 for the 5 replies. GLM 5.3 Flash's replies ran 133% longer, in tokens of reply, thinking not counted.

  • Agents and tool use: Grok 4.7 passed 5 of 5 and GLM 5.3 5 of 5. GLM 5.3 answered 2.8× sooner at the median, 2.5 s against 0.9 s. GLM 5.3 cost 5.8× less, $0.0129 against $0.0022 for the 5 replies. Their replies ran to about the same length.

The 5 prompts only one of Grok and GLM passed

Where one family's pick passed a prompt and the other's didn't, in each check's own words.

  • Announce a second bakery shop on LinkedIn (writing): GLM 5.3 passed and Grok 4.7 didn't. Grok 4.7: Graded 4.5 of 5 on average (lowest 4); but 94 words, under the 100 asked for. GLM 5.3: Graded 4.3 of 5 on average (lowest 4).

  • Rewrite corporate jargon in plain words (writing): GLM 5.3 passed and Grok 4.7 didn't. Grok 4.7: Graded 3.7 of 5 on average (lowest 3). GLM 5.3: Graded 4.0 of 5 on average (lowest 3).

  • Argue both sides of free buses (writing): Grok 4.7 passed and GLM 5.3 didn't. Grok 4.7: Graded 4.3 of 5 on average (lowest 4). GLM 5.3: Graded 3.7 of 5 on average (lowest 3); but a paragraph of 92 words, over the 90 allowed.

  • An article in three bullets (summarization): Grok 4.7 passed and GLM 5.3 didn't. Grok 4.7: Graded 4.7 of 5 on average (lowest 4). GLM 5.3: Graded 4.3 of 5 on average (lowest 3); but 68 words, over the 60 allowed.

  • A frustrated customer (customer support): GLM 5.3 passed and Grok 4.7 didn't. Grok 4.7: Graded 3.3 of 5 on average (lowest 3). GLM 5.3: Graded 4.7 of 5 on average (lowest 4).

Each Grok model against each GLM model

Every Grok model against every GLM model on the same 50 prompts: each one's passes, how many prompts split them, and who was cheaper and quicker.

  • Grok 4.7 vs GLM 5.3: 47 and 48 of 50; 5 prompts split them; GLM 5.3's replies cost 8.7× less in all, and GLM 5.3 answered sooner on 50, Grok 4.7 on 0. Grok 4.7 vs GLM 5.3.

  • Grok 4.7 vs GLM 5.3 Flash: 47 and 40 of 50; 9 prompts split them; GLM 5.3 Flash's replies cost 32.8× less in all, and GLM 5.3 Flash answered sooner on 31, Grok 4.7 on 15. Grok 4.7 vs GLM 5.3 Flash.

One writing prompt, both replies

Grok 4.7 and GLM 5.3 on a writing prompt, the job where they differed most: both replies as they came.

An everyday prompt: “Announce a second bakery shop on LinkedIn”, in full.

  • Grok 4.7

    Failed: Graded 4.5 of 5 on average (lowest 4); but 94 words, under the 100 asked for.

    We're opening a second Crumb & Co. shop in Leeds on 14 November.
    
    The new bakery will offer the same bread, cakes and pastries you already know from our first shop, baked fresh each morning. If you've been asking for a Leeds location, this one is for you.
    
    On opening day, coffee is free with any purchase. Come in, say hello, and try something from the counter. We'd love to meet you and hear what you'd like to see on the shelves.
    
    Thank you for supporting a small bakery. We can't wait to welcome you.

    1,624 tokens in, 207 out (93 of them reasoning) · 4.0 s · $0.0022 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  • GLM 5.3

    Passed: Graded 4.3 of 5 on average (lowest 4).

    Some news we've been itching to share: Crumb & Co. is opening a second shop!
    
    On 14 November, we'll be firing up the ovens in Leeds, bringing the same breads, pastries and cakes you've come to know us for to a brand-new corner of the world. It's a big step for a small bakery, and honestly, we couldn't be more excited (and a little nervous, in the best way).
    
    The new shop will have everything you love from our original spot — slow-fermented sourdough, buttery croissants, and weekend treats that disappear fast — plus a few surprises we're still perfecting.
    
    To celebrate opening day, every customer gets a free coffee with their purchase. So come along, say hello, and grab a cup on us.
    
    Leeds, we can't wait to meet you. See you on 14 November — come hungry.

    402 tokens in, 188 out · 2.9 s · $0.0004 · 1 message on Pro · answered by z-ai/glm-5.3 via Baidu ·

Every model, every job

All 3 Grok and GLM models in llmwise across the 50 prompts, with what a reply cost to run and each model's count on Pro.

Every Grok and GLM model across our test runs
ModelPassedHard onesCost per replyOn Pro
Grok 4.7xAI47 of 5020 of 20$0.0063Up to 250 a month
GLM 5.3Z.ai48 of 5019 of 20$0.00073Up to 250 a month
GLM 5.3 FlashZ.ai40 of 5016 of 20$0.00019Up to 60 a day

Passed: replies that passed their check, of those scored. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token. Every prompt and how it's scored.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

Its own subscription

The one company here with its own plan at about llmwise Pro's price ($20 a month), in its own words, dated. We drop a plan here when its facts are more than 45 days old.

  • SuperGrok (xAI), $30 a month

    Models it names: Grok 4.6. On its limits: “Smarter answers in Expert mode: Access our best model with higher limits” SuperGrok vs llmwise.

    Source: Grok plans, checked .

llmwise Pro, $20 a month, has all 3 of these models in one chat, from one monthly allowance: Grok 4.7 up to 250 messages a month and GLM 5.3 up to 250.

The lineups at a glance

What follows from each model's facts in our catalog.

  • The lineups

    Grok: one model, Grok 4.7. GLM: 2 models, GLM 5.3 and GLM 5.3 Flash.

  • Price per message

    Grok's one model is Grok 4.7 (250 messages a month on Pro); the least expensive GLM model is GLM 5.3 Flash (60 messages a day on Pro).

  • Context window

    Grok goes up to 500K tokens (Grok 4.7); GLM up to 1.05M tokens (GLM 5.3).

  • Images and PDFs

    GLM 5.3 doesn't read images. GLM 5.3 and GLM 5.3 Flash get a PDF's text rather than the file itself.

  • On the Free plan

    Free's one-time trial of 5 messages covers Grok 4.7, GLM 5.3, and GLM 5.3 Flash. Paid plans have every model, with messages every month.

Model by model

Two named models side by side, prompt by prompt, each with its messages on every plan.

More head-to-heads

Each of Grok and GLM against the other families, every job from the same test runs.

Where your messages go

Grok and GLM models are served only through OpenRouter, by endpoints that don't store or train on prompts. The Privacy Policy has the details.

Questions

Which is better, Grok or GLM?

In our test runs on September 27, 2026, the same 50 prompts across 10 jobs: Grok's one model passed 47 of 50 replies and GLM's 2 models passed 88 of 100 replies. Job by job, both families' best models shared the top result on 7 of the 10 jobs; Grok's alone had it on summarization; GLM's alone had it on writing and customer support.

Which is cheaper, Grok or GLM?

In llmwise, the least expensive Grok model is Grok 4.7 (250 messages a month on Pro), and the least expensive GLM model is GLM 5.3 Flash (60 messages a day on Pro). At API list prices (September 2026), a typical message of 4,000 tokens in and 700 out costs $0.0098 on Grok 4.7 and $0.0010 on GLM 5.3 Flash.

Can I use Grok and GLM in the same chat?

Yes. Pick a model for each message; when you switch, the next model sees the whole conversation, including the other one's answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.