Skip to content

Help

Limits and reliability

Every limit in llmwise in one place, what happens when you reach each one, and what llmwise does when a model's provider has trouble.

Limits

Message counts are on the pricing page; these are the rest.

  • Replies an hour

    Every plan has a soft cap of 100 replies in a rolling hour, counting new messages and approved tool steps together. At the cap, wait a little: the oldest replies drop out of the hour.

  • Reply length

    A reply can be up to 8k tokens on Free (Pro 8k, Max 16k, Ultra 16k, Studio 16k). At the limit it stops and offers Continue, which is a new message at the same price.

  • Long chats

    Once the conversation sent to the model passes 64k tokens, each reply counts as 2; past 128k, as 4; past 200k the chat is full. The message box says so before you send: “Long chat — each reply counts as 2. Start a new chat to save messages.”

  • Web searches

    On Free, web search works inside your 5 trial messages, up to 3 searches a day; reading a link counts as one. On a paid plan each search counts as one more message of the model in use. One message can search up to 5 times.

  • Files

    Up to 10 files a message, each up to 20 MB; images up to 5 MB, PDFs up to 100 pages.

  • Code and apps

    A code run stops after 60 seconds. A backend app runs 10 minutes, longer while you watch it, up to 30. Both are on paid plans.

  • Tool approvals

    One message can go back and forth with approved tools up to 10 times; after that, send a new message.

  • Skills

    Up to 20 skills of your own on Free, 100 on a paid plan; the built-in ones don't count.

  • Models on Free

    Free's one-time trial of 5 messages (one of them can be on Claude Opus 5.5) works on every model but Claude Fable 5.1 and GPT-6 Astra, which need a paid plan.

  • Fair-use limit

    On a paid plan: Pro $7.50, Max $20, Ultra $42 and Studio $85 of AI cost a billing period, at list prices; a $10 top-up adds $4 of room in the period it's bought in. At 80% we tell you; at the limit everything pauses until the period renews or you top up. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster.

When a count runs out

  • Out of everyday messages (a paid plan): the everyday models wait until the count refills at 00:00 UTC. The other models draw on the monthly allowance, so they keep working while it has messages left.

  • Out of monthly messages (a paid plan): the models that draw on the allowance stop until it renews or you top up. Everyday models keep working.

  • At the fair-use limit (a paid plan): every message pauses, the everyday models' too, until the period renews or you top up. We tell you at 80% first.

  • Out of Free's trial: Free is a one-time trial of 5 messages that doesn't refill, so the next message needs a paid plan.

How the counts work: plans, messages and limits.

When a provider has trouble

  • Claude, GPT, and Gemini go to the model's maker whenever llmwise can reach it directly. If the maker fails before the reply starts (an overload, a server error, a dropped connection), llmwise sends the same request to the same model through OpenRouter instead.

  • DeepSeek, Grok, Kimi, GLM, and Mistral are served through OpenRouter, and only by hosts that don't store or train on prompts; where a model has several hosts, OpenRouter moves to another when one is down.

  • Once a reply has started, a failure stops it: send the message again, or pick another model in the same chat.

  • A reply cut off because the provider failed costs nothing, however much of it arrived.

More in LLM failover, explained.

Questions

Are there API rate limits?

No. llmwise is a chat app: there's no API, so there are no request quotas or rate-limit headers to handle. What applies is your plan's message counts, on the pricing page, and the limits on this page.

Do I pay for a reply that fails?

Not when the provider failed. A reply cut off by a network error, an outage, an overload or a timeout costs nothing, however much of it arrived. A request the model rejects (a chat too long for it, a refusal) is charged like any reply once something came back.

Will llmwise answer with a different model if mine is down?

No. llmwise never switches you to a model you didn't pick. It retries the same model another way, and choosing a different one is up to you.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.