Skip to content

Blog · September 24, 2026

Updated · By , AI-assisted.

LLM failover: keeping answers coming when a provider fails

AI providers fail by overloading, going down or timing out, and apps survive with retries, a second route to the same model, a different model, or circuit breakers. When a model's maker fails before a reply starts, llmwise retries the same model through OpenRouter. It never swaps your model, and a reply the provider broke isn't counted.

How AI providers fail

  • Overload

    A provider under heavy load turns requests away with a rate limit or an “overloaded” error. Anthropic's API errors page lists one for exactly this: 529, overloaded.

  • Outages

    The provider's servers fail and answer with an error, for a minute or for hours. Each big provider logs its incidents on a status page, such as OpenAI's and Anthropic's.

  • Network failures and timeouts

    The connection drops, or the provider takes too long to answer at all.

  • Request errors (not outages)

    The request itself is the problem: the chat is too long for the model, or the model refuses. Retrying elsewhere doesn't fix these.

Four failover patterns

  • Retries with backoff

    Try the same request again after a short wait, waiting longer each time. Cheap, and enough for brief blips.

  • Another route, same model

    Send the same request to the same model through another route: another host, another region, or a service in front of several providers.

  • Another model

    Fall back to a different model when the first one is down. It keeps you answering, but the answer may differ in quality, style and price.

  • Circuit breakers

    After several failures in a row, stop sending to a route for a while and use the fallback straight away, instead of waiting on each request.

Most apps combine them: a quick retry, then another route, with a circuit breaker so a dead route doesn't slow every request. Falling back to a different model is the most drastic step, because the person asking gets an answer from a model they didn't choose.

How many routes does one model have?

Our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026 sent the same prompts to 19 models through OpenRouter, which reports the host that answered each one. 16 of the 19 came from a single host every time; the table names it. The open-weight models were another matter: 3 of them were answered by three hosts or more, and GLM 5.3 Flash by 9.

That's what a second route looks like in practice. A model many hosts serve can fail over to another host without anyone noticing. A model one company serves can only be retried, or reached through a cloud that also hosts it.

The hosts that answered our test requests, by model
ModelHostsMost replies from
Claude Fable 5.11Anthropic (50)
Claude Opus 5.51Claude Platform on AWS (50)
Claude Sonnet 5.51Anthropic (50)
Claude Sonnet 51Claude Platform on AWS (50)
Claude Haiku 5.51Anthropic (50)
Claude Haiku 4.51Amazon Bedrock (50)
GPT-6 Astra1OpenAI (50)
GPT-6.1 Sol1OpenAI (50)
GPT-6 Sol1OpenAI (50)
GPT-6 Luna1OpenAI (50)
Gemini 3.1 Pro1Google (50)
Gemini 3.8 Flash1Google (50)
DeepSeek V4.1 Flash1Together (50)
DeepSeek V4 Pro1Wafer (50)
Grok 4.71xAI (50)
Kimi K36Together (16), Wafer (13)
GLM 5.37Baidu (22), Wafer (21)
GLM 5.3 Flash9Wafer (21), AtlasCloud (15)
Mistral Large 41Mistral (50)

Hosts as OpenRouter named them in each reply. The runs, prompt by prompt.

What llmwise does

  • Claude, GPT, and Gemini: sent to the model's maker whenever llmwise can reach it directly. If the maker fails before the reply starts (an overload, a server error, a dropped connection), llmwise sends the same request to the same model through OpenRouter instead.

  • DeepSeek, Grok, Kimi, GLM, and Mistral: served only through OpenRouter, by hosts that don't store or train on prompts. Where a model has more than one such host, OpenRouter moves the request to another when one is down.

  • Once a reply has started, a failure isn't retried behind your back: the reply stops, and you can send it again or switch models.

  • It never swaps your model for another. If a model keeps failing, pick a different one in the model picker: it sees the whole chat.

What a failure costs you

Nothing, when the provider is at fault: a reply cut off by a network error, an outage, an overload or a timeout isn't counted, however much of it had arrived. A request the model itself rejects is different: it's counted like any reply once something came back, and not counted if nothing did. The limits and reliability page has the details.

Bar chart: How many hosts answered each model. GLM 5.3 Flash: 9 hosts; GLM 5.3: 7 hosts; Kimi K3: 6 hosts; Claude Fable 5.1: 1 host; Claude Opus 5.5: 1 host; Claude Sonnet 5.5: 1 host; Claude Sonnet 5: 1 host; Claude Haiku 5.5: 1 host; Claude Haiku 4.5: 1 host; GPT-6 Astra: 1 host; GPT-6.1 Sol: 1 host; GPT-6 Sol: 1 host; GPT-6 Luna: 1 host; Gemini 3.1 Pro: 1 host; Gemini 3.8 Flash: 1 host; DeepSeek V4.1 Flash: 1 host; DeepSeek V4 Pro: 1 host; Grok 4.7: 1 host; Mistral Large 4: 1 host.
The hosts OpenRouter reported for each reply in our test runs of September 27, 28, 29 and October 2, 7, 8, 9, 2026.

Failover: questions

Does llmwise switch me to another model when one fails?

No. llmwise never swaps the model you picked for a different one. When a route fails it tries the same model another way; choosing a different model is up to you.

Am I charged for a reply that failed?

Not when the provider failed. A reply cut off by a network error, an outage, an overload or a timeout costs nothing, however much of it had arrived. A request the model rejects, such as a chat too long for it, is charged like any other reply once something came back.

When one model's provider has a bad day

Switch to another of llmwise's 19 models in the middle of a chat; it picks up the whole conversation.