Skip to content
DineshKumar Sarangapani
All writing

Resilient Multi-Cloud LLM Gateways: Routing, Rate Limits, and Failovers

How to design a production model gateway that handles provider outages, rate limits, and latency spikes across multiple cloud providers.

If your enterprise AI application connects directly to a single cloud provider’s LLM endpoint, your application will eventually experience downtime.

All foundation model APIs suffer from operational challenges:

  • Global capacity shortages during peak hours.
  • Sudden rate-limit errors when token bucket quotas are exhausted.
  • Regional latency spikes where Time To First Token jumps from 400 milliseconds to several seconds.
  • Partial brownouts where streaming connections hang indefinitely without emitting bytes.

When critical internal systems and production users rely on AI agents, a cloud provider outage cannot be allowed to halt business operations.

To solve this, we architected a resilient multi-cloud model gateway. It presents a standard OpenAI-compatible interface to internal applications while dynamically orchestrating requests across multiple cloud providers.

Here is the idea, how it performed in production, and what to watch out for.


The Idea: A Resilient Model Proxy

Rather than letting each internal application talk directly to external LLM vendors, all traffic routes through a centralized gateway service.

The gateway manages three core responsibilities:

  1. Dynamic Provider Routing: Applications request a logical model capability (such as general-purpose completion or complex reasoning). The gateway maps that capability to healthy providers and regions based on priority, latency, and cost.
  2. Circuit Breakers and Fast Failover: If a primary cloud provider returns errors or times out, the gateway immediately fails over to an alternative provider or region before the user experiences an error.
  3. Upstream Rate Limit Coordination: The gateway tracks provider capacity and respects throttling signals so that traffic spikes do not cause cascading failures.

How It Worked Well

  1. High Availability via Automated Failover: When one cloud region experienced degraded capacity, the gateway automatically diverted subsequent requests to healthy regions or secondary cloud providers. End users experienced a brief delay of a few hundred milliseconds rather than a failed run.
  2. Standardized Developer Experience: Application teams developed against one consistent API contract. They did not need to write custom SDK code for each cloud provider, and changing underlying models required zero application code modifications.
  3. Throttling Protection with Rate-Limit Forwarding: When an upstream provider returned an HTTP 429 response, the gateway captured the upstream wait window and forwarded it to clients while temporarily tripping a local circuit breaker. This prevented retry storms from overwhelming overloaded providers.
  4. Actionable Latency Observability: Instead of tracking total request time, the gateway isolated Time To First Token (TTFT) from Time Per Output Token (TPOT). This allowed on-call engineers to distinguish between upstream scheduling queues and slow token generation throughput.

What to Watch Out For

  1. Prompt Replay on Streaming Failures: If an upstream provider fails before emitting any tokens, the gateway can cleanly replay the prompt to a secondary provider. However, if the stream fails mid-generation after emitting 50 tokens, replaying the entire prompt from scratch will duplicate text in the user’s chat. The gateway must detect whether bytes have traversed to the client and handle partial stream interruptions gracefully.
  2. Behavioral Divergence Across Models: Even when two models are considered comparable, their system prompt interpretation, tool calling formats, and output schemas may subtly differ. Ensure your agents are tested against the fallback model families so that an automated failover does not produce unexpected tool argument syntax.
  3. Streaming Deadlocks and Connection Leaks: When users abandon a web page or cancel a request mid-stream, client connections close. If your gateway does not actively propagate cancellation signals upstream, the underlying cloud connection remains open, continuing to consume model tokens and holding sockets in your connection pool.
  4. Circuit Breaker Flapping: If your circuit breaker threshold is too aggressive, a few transient network drops will unnecessarily route all traffic to more expensive backup providers. Use a combination of consecutive failure counts and cooldown intervals with canary probe requests before fully reopening circuits.