Engineering14 min read

LLM Gateway for AI Agents: Routing, Failover and Cost Control

Learn how an LLM gateway routes AI agent requests, handles provider failures and rate limits, and controls spend across models.

OpenAI documents separate request and token rate limits, because one model call can consume far more capacity than another. For an AI agent that may make dozens of dependent calls, that distinction is the difference between a controlled workload and a failed run.

An LLM gateway sits between an agent and its model providers. It authenticates each request, chooses an eligible model, applies rate and budget policy, and handles approved retries or fallbacks. It can make model access more resilient, but it cannot make every model interchangeable.

An LLM gateway should be a policy enforcement point, not merely a unified API proxy.

This guide explains how to design that layer for routing, failover, rate limits, and spend control. It deliberately stays below agent orchestration and above model providers.

Connect models and tools to an AI agent running on Rerun

What is an LLM gateway?

An LLM gateway is a model-aware intermediary between software and one or more inference providers. The agent sends a standardized request to the gateway. The gateway identifies the workload, checks policy, selects a target, sends the request, records usage, and returns the result.

A simple path looks like this:

AI agent -> LLM gateway -> provider or model A
                        -> provider or model B
                        -> self-hosted model

The unified endpoint is convenient. Centralized policy is the larger benefit. Without it, every agent can accumulate its own provider SDKs, keys, retry loops, quota rules, and cost calculations.

The request path

A production request usually follows eight steps:

  1. Authenticate the application, tenant, and agent.
  2. Read capability requirements and request metadata.
  3. Verify budgets and rate limits.
  4. Remove unhealthy or ineligible deployments.
  5. Rank the remaining targets.
  6. Reserve capacity and estimated spend.
  7. Send the request and apply a bounded failure policy.
  8. Reconcile actual usage and return routing metadata.

The identity at step one matters. If all calls look identical, the gateway cannot stop one looping agent or noisy tenant from consuming everybody's capacity.

LLM gateway vs API gateway, router, and agent framework

LayerPrimary responsibilityWhat it does not own
Traditional API gatewayGeneral authentication, traffic, and API policiesModel capabilities and token economics
LLM gatewayModel-aware routing, provider fallback, token limits, and inference spendThe complete agent job
Model routerSelects a model or deploymentBroader authentication and budget enforcement, unless bundled
Agent frameworkPlans steps, calls tools, and manages stateCentral model-provider policy
Agent operations layerRuns, schedules, isolates, and supervises agentsPer-request model routing

Vendor language overlaps. "AI gateway," "LLM proxy," and "model gateway" may describe similar products, but a shared endpoint alone does not guarantee capability routing, hard budgets, or safe fallbacks.

Why AI agents need a model-access layer

A chatbot request is often one exchange. An agent task can be a chain of planning, extraction, tool selection, verification, and final-response calls. Parallel subtasks create bursts. Long-running agents are more likely to encounter a provider incident. A failure near the end can waste tokens and tool activity already consumed.

Direct provider integrations also spread policy through the codebase. One service retries three times, another retries forever, and a third silently selects a cheaper model. Centralizing those choices makes them reviewable.

This is where flowchart automation falls short. Zapier, Make, and n8n are useful for deterministic handoffs, but a box connected to another box does not define safe model substitution. A chatbot also cannot enforce a fleet-wide token budget simply by answering in a window.

When you do not need a gateway

A gateway is not mandatory for every prototype. Direct provider access may be simpler when:

  • one provider and one model handle the workload;
  • traffic and spending are low and predictable;
  • provider-specific features are essential;
  • centralized quotas and failover are unnecessary;
  • another network hop would create more risk than it removes.

Use a gateway when the policy has become real, not because the category is fashionable.

Model routing: choose eligibility before optimization

The safest routing rule is simple: first determine which models can complete the request correctly, then optimize for cost or speed.

Static routing

Static routing maps a stable alias to a deployment. Agent code can request fast-model, reasoning-model, or fallback-model without embedding a provider name everywhere. Mappings can differ by environment, customer plan, region, or workload.

Aliases reduce coupling, but they also create responsibility. Changing an alias can change output behavior without changing application code. Treat alias updates like production releases.

Capability-based routing

An eligibility filter can require:

  • reliable tool or function calling;
  • structured JSON output;
  • vision or audio support;
  • a minimum context window;
  • regional availability;
  • approved retention and data-processing terms;
  • an allowed provider list;
  • a latency class;
  • a validated output-quality threshold.

A model that fails one hard requirement should not enter the optimization pool, however cheap it is.

Cost-, latency-, and quality-aware routing

Once the eligible pool is known, rank it using current signals. Useful inputs include expected input and output tokens, price, time to first token, end-to-end latency, recent errors, remaining quota, and task-specific evaluation scores.

Agent stepFirst routing prioritySecondary constraint
ClassificationStructured-output reliabilityLow cost
PlanningReasoning qualityContext capacity
Tool selectionFunction-calling accuracyLatency
SummarizationQuality thresholdCost and speed
Customer responseQuality and toneResponse time
Background enrichmentThroughputQuota headroom

A global "cheapest model" policy is usually a false economy. A weak planning call can cause extra tool calls, retries, and corrections that cost more than the model saving.

Models are not automatically interchangeable

Compatible APIs hide meaningful differences. Models may interpret tool schemas differently, produce different JSON shapes, count tokens differently, reject parameters, or apply different safety behavior. Versions behind an alias may also change.

Before a model joins a fallback pool, test it on the actual task's golden dataset. Measure schema validity, tool-choice accuracy, task outcome, latency, and cost. An endpoint that accepts the same request is not proof that it preserves the same semantics.

Router - Load Balancing | liteLLMdocs.litellm.ai

Failover and retries without making failures worse

Fallback policy should distinguish three actions:

  1. retry the same deployment for a short transient failure;
  2. switch to an equivalent deployment when provider capacity is unhealthy;
  3. switch models only when the task has validated substitutes.

Provider guidance matters. AWS documents retry modes that use exponential backoff, jitter, and token buckets. The goal is not to retry everything. It is to spend a small, explicit retry budget on failures that can plausibly recover.

Which failures should trigger fallback?

FailureTypical responseReason
429 rate limitRespect retry guidance, queue briefly, or route to approved capacityThe request may be valid
TimeoutRetry within the deadline or use a healthy equivalent targetThe outcome may be unknown
Connection errorBack off and use another healthy deploymentThe endpoint may be unavailable
Provider 5xxBounded retry, circuit breaker, then fallbackUsually transient
Invalid requestDo not retry blindlyThe payload must change
Authentication errorStop and repair credentialsRepetition will not help
Context too largeSelect a validated larger-context route or change the requestSame-target retry is pointless

Set maximum attempts, a total retry deadline, exponential backoff with random jitter, and per-run retry budgets. Respect provider headers. Open the circuit for a repeatedly failing deployment so every agent does not rediscover the same outage.

Streaming changes the failure boundary

Failover is relatively clean before any output reaches a user. Once tokens have streamed, another model may restart, duplicate, or contradict the partial answer. Choose a policy explicitly:

  • buffer output until a safe boundary;
  • restart visibly;
  • return a labeled partial result;
  • fail the step and let the caller decide.

There is no universal choice. Interactive support and background extraction have different latency and consistency needs.

Protect tool side effects

Retrying a model call is not the same as retrying an entire agent step. If a previous attempt already sent an email, updated a CRM, or issued a refund, replaying the step can duplicate the action.

Give tool operations idempotency keys and store their completion state outside the model response. Keep gateway retries on the model-call side of that boundary. Broader recovery belongs in the AI agent infrastructure layer.

Monitor every action taken by an always-on AI agent

Rate limiting for agent traffic

Requests per minute are not enough because request sizes vary by orders of magnitude. Anthropic's rate-limit documentation and OpenAI's guidance both show why provider quotas need more than one dimension.

Track at least:

  • requests per minute, or RPM;
  • tokens per minute, or TPM;
  • concurrent requests;
  • daily or monthly token allowances;
  • spend per time window;
  • provider-specific quota headroom.

Apply limits hierarchically across the organization, environment, team, tenant, agent, API key, and model. A child scope should never be able to bypass its parent's ceiling.

Queue, reject, degrade, or reroute?

ConditionReasonable response
Short traffic spikeQueue within a strict latency window
Provider quota reachedUse an equivalent approved deployment
Tenant allowance exceededReject with a clear limit response
Premium-model budget depletedRoute only to a validated lower-cost model
Interactive request near deadlinePrefer a fast fallback over a long retry
Background jobDelay or reschedule instead of paying a premium

Do not degrade silently. Return routing metadata, record which model handled the call, and surface material quality changes when they affect the user.

Spend controls that stop runaway agent costs

A dashboard shows what happened. It does not stop the next expensive request. Enforcement must sit in the request path.

Before sending, estimate input tokens and reserve expected output cost. Account for the chosen provider, current price metadata, caching rules, and reasoning tokens where relevant. After completion, reconcile estimated usage with actual usage.

Hard caps, soft alerts, and budget hierarchy

Useful controls include:

  • maximum output tokens and estimated cost per request;
  • per-agent and per-tenant budgets;
  • daily and monthly caps;
  • warning thresholds;
  • hard stops;
  • approval for exceptional requests;
  • inheritance rules when scopes overlap.

For example, an agent may have $100 remaining while its tenant has $0. The stricter parent budget must win.

Cost-aware routing without quality collapse

Use this sequence:

  1. define a minimum quality threshold per task;
  2. admit only models that pass task-level evaluation;
  3. choose the least expensive eligible option;
  4. escalate complex or low-confidence cases;
  5. reevaluate when models, prices, or prompts change.

For total cost and ROI, use the dedicated AI agent cost guide. The gateway's narrower job is to enforce request-level limits before a runaway loop becomes a bill.

A reference routing and fallback policy

The following agent brief is illustrative. It is not drop-in production configuration, but it makes the policy order concrete.

LLM gateway policy for a support agent
{
  "workload": "support-agent",
  "requirements": {
    "tool_calling": true,
    "structured_output": true,
    "minimum_quality_score": 0.9
  },
  "routing": {
    "objective": "lowest_cost_within_quality_threshold",
    "allowed_regions": ["us", "eu"]
  },
  "limits": {
    "rpm": 120,
    "tpm": 200000,
    "max_estimated_cost_per_request_usd": 0.25
  },
  "retries": {
    "retry_on": [429, 500, 502, 503, 504, "timeout"],
    "max_attempts": 3,
    "backoff": "exponential_with_jitter"
  },
  "fallbacks": [
    "equivalent-provider-deployment",
    "validated-secondary-model"
  ]
}

The execution order should be deterministic: validate identity and budget, determine capabilities, remove unhealthy targets, rank the eligible pool, reserve spend, send the request, retry within the deadline, then reconcile usage.

How to evaluate an LLM gateway

Do not choose from a feature checklist alone. Test the behavior under the failures your agents will actually face.

Managed gateways reduce operational work but add a vendor and network dependency. Self-hosted gateways offer control but require patching, scaling, and on-call ownership. Cloud-platform gateways may integrate well with existing identity and networking but increase platform coupling. Direct provider access remains a valid baseline for simpler systems.

LLM gateway observability: what to log and measure

A gateway can explain request-level behavior only if it records the routing decision, not merely the final model. Give every call a request ID and connect it to the agent run's trace ID. Log the agent, tenant, task, provider, model, deployment, rejected candidates, retry chain, and fallback reason.

Measure input, output, cached, and reasoning tokens separately when the provider exposes them. Record estimated versus actual cost, time to first token, total latency, error class, and every budget or rate-limit decision. Store prompt and policy versions, but redact secrets and sensitive payload data.

A useful production trace should answer a concrete question such as: "Why did this support task use the secondary model?" The answer might be: primary deployment rejected by TPM limit, equivalent deployment returned 429, validated fallback selected within the remaining latency and cost budget.

That record is an operational artifact, not just a dashboard event. During incident review, the team should be able to reconstruct this sequence:

Trace fieldExample valueWhy it matters
Primary targetreasoning-model-usShows the intended route
RejectionTPM headroom below request estimateExplains why it was skipped
Second targetreasoning-model-euPreserves capability
Failure429 after 240 msSeparates quota from model quality
Fallbackvalidated-secondary-modelShows the approved substitute
Final resultsuccess, $0.08, 1.9 sReconciles outcome, cost, and latency

This is deliberately narrower than full AI agent observability. Gateway telemetry follows model requests. Agent operations must connect those calls to tool actions, approvals, state changes, and the outcome of the complete job.

Where the LLM gateway fits in an AI agent stack

The gateway sits below agent logic and above model providers. It governs individual model requests. It does not own the life of a multi-step job, schedule recurring work, manage tool side effects, or provide the whole execution environment.

The gateway governs each model request. The agent operations layer keeps the broader agent process running.

Tool typeBest forLimitation for operating agents
Zapier, Make, and n8nDeterministic app-to-app workflowsEvery branch must be predefined
ChatbotsUser-led conversationsThey wait for input rather than supervise ongoing work
LLM gatewayGoverning model requestsIt does not run the complete agent lifecycle
Agent operations layerPersistent agent executionIt complements rather than replaces the gateway

Rerun belongs to that operations layer, not the gateway layer. It gives ready-to-run or custom agents an always-on private cloud machine, connected tools, visible logs, schedules, and human approvals. A Rerun agent can depend on a third-party or self-hosted LLM gateway without confusing the two responsibilities.

That division also exposes the weakness of chat windows and flowcharts. A chat interface waits for a person to return. A deterministic automation canvas requires someone to wire every branch. An operating agent must keep working, pause safely for judgment, and make its actions inspectable.

For the complete runtime boundary, read AI Agent Infrastructure. For deployment ownership, see Self-Hosted AI Agents.

AI Agent Infrastructure: The Runtime & Serving Stack

AI Agent Infrastructure: The Runtime & Serving Stack

AI agent infrastructure is the runtime and serving stack that keeps agents safe in production. The layers, the control plane, and why 40% of projects fail.

The practical rule

An LLM gateway is worth adopting when model access has become a shared production dependency. Design it as an enforcement layer. Filter for capability first, optimize second, cap retries, isolate tenants, reserve spend, and test every fallback on real tasks.

It will not make models identical. It will not replace durable agent execution. It will not turn a Zap, a flowchart, or a chatbot into an autonomous operating system.

Use the gateway to govern model access. Use an operations layer to run the agent that depends on it.

Frequently asked questions

What is an LLM gateway?

An LLM gateway is a model-aware layer between an application or AI agent and model providers. It authenticates requests, applies routing and budget policies, manages rate limits, and handles approved retries or fallbacks.

Does an LLM gateway reduce LLM costs?

It can. A gateway enables cost-aware routing, hard budgets, token limits, and centralized usage accounting. Savings depend on workload design, model eligibility rules, and whether cheaper models meet the required quality threshold.

What is the difference between an LLM gateway and an API gateway?

A traditional API gateway manages general authentication and traffic policies. An LLM gateway adds model-aware features such as token accounting, capability routing, provider fallback, inference budgets, and model-specific rate limits.

Can an LLM gateway automatically switch models?

Yes, but automatic switching is safe only when the alternate model has been validated for the task. Compatible endpoints do not guarantee identical tool calls, structured outputs, context limits, or safety behavior.

How does an LLM gateway handle provider rate limits?

It can track RPM, TPM, concurrency, and quota headroom, then queue briefly, back off, route to approved capacity, or reject the request according to policy.

Is an LLM gateway a single point of failure?

It can be. Production deployments need redundancy, health checks, load balancing, circuit breakers, and a clearly defined bypass or failure policy.

Do AI agents need an LLM gateway?

Not always. A gateway becomes valuable when agents use multiple providers, share quotas, require production failover, or create enough inference spend to justify centralized enforcement.

What is the difference between an LLM gateway and a model router?

A model router selects a model or deployment. An LLM gateway may include routing, but also centralizes authentication, rate limits, retries, provider fallback, usage accounting, and budget enforcement.

What is the difference between an LLM gateway and an agent framework?

An LLM gateway governs model requests. An agent framework defines planning, tools, state, and orchestration. They solve different layers and can be used together.

Clément Janssens

Written by

Clément Janssens

Related articles

Your first agent is
three minutes away

Start for free

© 2026 Rerun. All rights reserved.