One LLM gateway for every provider, ready in 3.2 milliseconds
The open-source LLM gateway behind Swarms. Call OpenAI, Anthropic, Gemini, Groq, OpenRouter, Ollama and any OpenAI-compatible server through one function, with LiteLLM's function names and 383x faster imports.
v0.2.0Apache 2.0Python 3.10+Paper
import routehub
messages = [{"role": "user", "content": "Summarize the Q3 risks."}]
for model in [
"gpt-5.4-mini",
"claude-sonnet-4-6",
"gemini/gemini-3-flash-preview",
"groq/llama-3.3-70b-versatile",
]:
response = routehub.completion(model=model, messages=messages)
print(response.choices[0].message.content)Change the model string to change providers. The call and the response type stay the same.
Benchmarks
383x faster to import than LiteLLM, and 5.5x faster to a first response
Every library ran in its own environment against an instant mock server, so the numbers isolate what the client itself costs. The bare provider SDK is the floor: the same HTTP and parsing work with no gateway in front of it.
- to import
- 3.2 ms
- fresh process to first response
- 267 ms
- peak memory
- 54 MiB
- added to each OpenAI call
- 9 µs
LiteLLM takes 1,235 ms, 383x longer
LiteLLM takes 1,464 ms, 5.5x longer
LiteLLM uses 211 MiB, 3.9x more
LiteLLM adds 647 µs
Time to import the library
Median of 20 fresh processes, in milliseconds
383x faster
- RouteHub3.2 ms
- OpenAI SDK (bare)68x217.2 ms
- LiteLLM386x1,235.3 ms
Importing RouteHub loads 18 modules and no third-party code. LiteLLM loads 2,502, including the OpenAI SDK, aiohttp, jinja2 and tokenizers, and with its default settings it also downloads a model price list from GitHub before your code runs.
Fresh process: import plus the first request
Median of 20 fresh processes, in milliseconds
5.5x faster
- RouteHub267 ms
- OpenAI SDK (bare)258 ms
- LiteLLM5.5x1,464 ms
RouteHub loads the OpenAI SDK when the first request needs it, so import plus first call is the fair comparison. It lands within 9 ms of the bare SDK. Serverless functions, CLI tools and CI jobs pay this on every start.
Peak memory after the first request
Peak resident set size, in MiB
3.9x less
- RouteHub54 MiB
- OpenAI SDK (bare)53 MiB
- LiteLLM3.9x211 MiB
RouteHub stays within about a megabyte of the bare OpenAI SDK. LiteLLM's process holds 194 MiB before it sends a single request.
Latency of a warm call
Median of 1,200 sequential calls, in milliseconds
+9 µs
- RouteHub0.55 ms
- OpenAI SDK (bare)0.54 ms
- LiteLLM2.2x1.19 ms
RouteHub adds 9 µs to the bare SDK, inside run-to-run noise. LiteLLM adds 647 µs to every call: it creates a logging object, runs cache and metadata hooks, converts the reply into its own type and hands a success handler to a thread pool.
Consuming a 200-chunk stream
Median of 200 streams, in milliseconds
6.7x faster
- RouteHub9.1 ms
- OpenAI SDK (bare)9.8 ms
- LiteLLM6.7x61.1 ms
LiteLLM spends about 224 µs on every chunk and delivers the first one after 16.6 ms. At 100 chunks per second, 50 open streams at that cost keep a full CPU core busy on chunk handling alone.
A call after a 10-second tool pause
Real requests to api.openai.com, median of 6 runs, in milliseconds
67 ms saved
- RouteHub116 ms
- OpenAI SDK (bare)1.4x166 ms
- LiteLLM1.6x185 ms
httpx and the OpenAI SDK close idle connections after 5 seconds. RouteHub keeps them open for 60, so an agent step that follows a long tool call skips the TCP and TLS handshake. Over a 50-step agent run that is about 3.3 seconds.
What a clean install pulls in
Fully resolved dependency tree
21 vs 58
- RouteHub21 packages
- LiteLLM2.8x58 packages
RouteHub has three direct dependencies: openai, pydantic and tiktoken. Every installed package is code that runs with your API keys, so a smaller tree is a smaller attack surface.
What it adds up to over an agent run
Describe your agent and see the client-side time each gateway adds, before the model generates a single token.
Runs per day
4.52 s
saved on every run
13 hours of agent wall-clock time a day at 10,000 runs.
| Where the time goes | LiteLLM | RouteHub |
|---|---|---|
| Fresh process to first response | 1.46 s | 267 ms |
| Gateway overhead across 50 calls | 37 ms | 1.6 ms |
| Reconnects after 49 tool pauses | 3.28 s | 0.0 ms |
RouteHub (commit 55284e1) against LiteLLM 1.104.0 on an Apple M3 Pro with CPython 3.12, each in its own environment. The calculator uses the 22-message agent request for per-call overhead and the 67 ms handshake measured against api.openai.com; network savings scale with your round-trip time to the provider. The paper also covers async throughput, where LiteLLM's aiohttp transport pulls ahead at 128 requests in flight, and results for any-llm and aisuite.
Full methodology in the paperQuickstart
Install it, set one key, and call any model
You need Python 3.10 or later and an API key for any supported provider. Each step builds on the one before it.
- 1
Install from PyPI
Python 3.10 or later. RouteHub has three direct dependencies: openai, pydantic and tiktoken. Add the fast extra to use orjson for JSON.
terminalpip install routehub # or with uv uv add routehub # with orjson for faster JSON handling pip install "routehub[fast]"terminal · example output - 2
Set a key and make your first call
Export the key for your provider, then call completion(). The response is the OpenAI SDK's own ChatCompletion, so attribute access and type hints work as usual.
main.py# export OPENAI_API_KEY="sk-..." import routehub response = routehub.completion( model="gpt-5.4-mini", messages=[{ "role": "user", "content": "Name three risks of a single-region deployment.", }], ) print(response.choices[0].message.content) print(response.usage.total_tokens, "tokens")terminal · example output - 3
Change providers by changing one string
Prefix the model with its provider, or use a bare name like claude-sonnet-4-6. Set the key for each provider you call. The function and the response type stay the same.
providers.pyimport routehub messages = [{"role": "user", "content": "Reply with one word: ready?"}] for model in [ "gpt-5.4-mini", "claude-sonnet-4-6", "gemini/gemini-3-flash-preview", "groq/llama-3.3-70b-versatile", ]: response = routehub.completion(model=model, messages=messages) print(f"{model:32} {response.choices[0].message.content}")terminal · example output
Features
Everything an LLM gateway needs
Pick a feature to see how it works and what it prints. The tabs above each sample switch between providers, approaches, or a before and after.
Call any model
Every major provider
OpenAI, Anthropic, Gemini, Groq, xAI, DeepSeek, OpenRouter, Together, Mistral, Fireworks, Azure OpenAI, Ollama, vLLM and any OpenAI-compatible server, all through one function.
- Requests use the OpenAI chat format on every provider
- Responses are the OpenAI SDK's own ChatCompletion, ChatCompletionChunk and CreateEmbeddingResponse
- Bare names starting with gpt-, claude-, gemini-, grok-, deepseek- or mistral- need no prefix
- api_base= points any model at a proxy, a private endpoint or a local server
import routehub
messages = [{"role": "user", "content": "Draft a release note."}]
routehub.completion(model="gpt-5.4", messages=messages)
routehub.completion(model="claude-sonnet-4-6", messages=messages)
routehub.completion(model="gemini/gemini-3-flash-preview", messages=messages)
routehub.completion(model="xai/grok-4", messages=messages)
routehub.completion(model="deepseek/deepseek-chat", messages=messages)
routehub.completion(
model="openrouter/anthropic/claude-sonnet-4.6",
messages=messages,
)import routehub
# Ollama on this machine, or set OLLAMA_API_BASE for a remote one
routehub.completion(model="ollama/llama3", messages=messages)
# A vLLM server
routehub.completion(
model="hosted_vllm/meta-llama/Llama-3.3-70B-Instruct",
messages=messages,
api_base="http://gpu-box:8000/v1",
)
# Any OpenAI-compatible server
routehub.completion(
model="support-model",
messages=messages,
api_base="https://llm.internal.example.com/v1",
)import os
import routehub
# AZURE_API_KEY, AZURE_API_BASE and AZURE_API_VERSION
# come from the environment...
routehub.completion(model="azure/gpt-prod", messages=messages)
# ...or from the call, so two tenants can use two resources
routehub.completion(
model="azure/gpt-prod",
messages=messages,
api_base="https://acme.openai.azure.com",
api_version="2024-10-21",
api_key=os.environ["ACME_AZURE_KEY"],
)Call any model
Streaming and async
Stream tokens as they arrive, or send requests concurrently with acompletion. Both take exactly the same arguments as completion.
- Streams yield ChatCompletionChunk objects and work in a with block
- stream_options={"include_usage": True} adds a final usage chunk, exactly as OpenAI does
- aembedding and async model listing are there too
- Async HTTP clients are cached per event loop
import routehub
stream = routehub.completion(
model="claude-sonnet-4-6",
messages=[{"role": "user", "content": "What is an idempotency key?"}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.usage:
print(f"\n{chunk.usage.total_tokens} tokens")
elif chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)import asyncio
import routehub
async def main():
questions = ["Japan", "Kenya", "Peru"]
responses = await asyncio.gather(*(
routehub.acompletion(
model="gpt-5.4-mini",
messages=[{
"role": "user",
"content": f"Capital of {country}? One word.",
}],
)
for country in questions
))
print([r.choices[0].message.content for r in responses])
asyncio.run(main())Call any model
Tool calling
Define tools once in the OpenAI format and use them on any provider. For Claude, RouteHub translates definitions, calls and results to Anthropic's format and back, including parallel calls.
- Read calls from message.tool_calls on every provider
- Return results as {"role": "tool"} messages with the call's id
- tool_choice is translated along with the tools
import json
import routehub
weather = {
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
model = "claude-sonnet-4-6"
messages = [{"role": "user", "content": "What's the weather in Paris?"}]
response = routehub.completion(model=model, messages=messages, tools=[weather])
message = response.choices[0].message
call = message.tool_calls[0]
print(call.function.name, json.loads(call.function.arguments))
messages += [
{"role": "assistant", "content": message.content,
"tool_calls": [call.model_dump()]},
{"role": "tool", "tool_call_id": call.id, "content": "Sunny, 21°C"},
]
final = routehub.completion(model=model, messages=messages, tools=[weather])
print(final.choices[0].message.content)import json
import routehub
weather = {
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
model = "gemini/gemini-3-flash-preview"
messages = [{"role": "user", "content": "What's the weather in Paris?"}]
response = routehub.completion(model=model, messages=messages, tools=[weather])
message = response.choices[0].message
call = message.tool_calls[0]
print(call.function.name, json.loads(call.function.arguments))
messages += [
{"role": "assistant", "content": message.content,
"tool_calls": [call.model_dump()]},
{"role": "tool", "tool_call_id": call.id, "content": "Sunny, 21°C"},
]
final = routehub.completion(model=model, messages=messages, tools=[weather])
print(final.choices[0].message.content)Call any model
Structured output
Pass a pydantic model and get back JSON that validates against it. For Claude, the schema is enforced through a forced tool call and returned as the message content.
- A pydantic model becomes a strict json_schema response format
- Validate the reply with Model.model_validate_json
- The same code works on OpenAI and Claude
from pydantic import BaseModel
import routehub
class RiskReport(BaseModel):
title: str
severity: int
response = routehub.completion(
model="gpt-5.4-mini",
messages=[{
"role": "user",
"content": "Assess the main risk of a single-region deployment.",
}],
response_format=RiskReport,
)
report = RiskReport.model_validate_json(
response.choices[0].message.content
)
print(report)from pydantic import BaseModel
import routehub
class RiskReport(BaseModel):
title: str
severity: int
response = routehub.completion(
model="claude-sonnet-4-6",
messages=[{
"role": "user",
"content": "Assess the main risk of a single-region deployment.",
}],
response_format=RiskReport,
)
report = RiskReport.model_validate_json(
response.choices[0].message.content
)
print(report)Claude, natively
Prompt caching and thinking
Anthropic's OpenAI-compatible endpoint drops prompt caching, thinking output and cached-token usage. RouteHub calls the native Messages API instead, so all three work.
- cache_control markers on messages and tools go to Anthropic unchanged
- reasoning_effort turns on extended thinking; it comes back as reasoning_content and signed thinking_blocks
- Cache reads and writes are counted in usage, so cost accounting stays correct
- Images, PDFs and tool use are translated both ways
import routehub
messages = [
{
"role": "system",
"content": [{
"type": "text",
"text": policy_handbook,
"cache_control": {"type": "ephemeral"},
}],
},
{"role": "user", "content": "Which policies cover vendor onboarding?"},
]
for attempt in (1, 2):
response = routehub.completion(
model="claude-sonnet-4-6", messages=messages
)
usage = response.usage
print(
f"call {attempt}:",
f"read {usage.prompt_tokens_details.cached_tokens}",
f"written {usage.cache_creation_input_tokens}",
)import routehub
response = routehub.completion(
model="claude-sonnet-4-6",
messages=[{"role": "user", "content": "Is 2027 a prime number?"}],
reasoning_effort="medium",
)
message = response.choices[0].message
print("thinking:", message.reasoning_content)
print("answer:", message.content)
# Send message.thinking_blocks back on the assistant turn
# for multi-turn tool use, as Anthropic requires.Run in production
Typed errors and retries
Every provider failure is raised as a typed exception that carries the status code, provider and model. Each one subclasses the matching OpenAI SDK exception, so the handlers you already have keep working.
- RateLimitError, ContextWindowExceededError, AuthenticationError, ServiceUnavailableError and more
- num_retries retries rate limits, 5xx errors and connection failures with exponential backoff that respects retry-after
- The original error is chained as __cause__
import routehub
try:
routehub.completion(
model="gpt-5.4-mini",
messages=messages,
num_retries=3,
)
except routehub.ContextWindowExceededError:
messages = messages[-10:] # trim the conversation and retry
except routehub.RateLimitError as error:
print(error.status_code, error.llm_provider, error.model)import routehub
try:
routehub.completion(
model="gpt-5.4-mini",
messages=[{"role": "user", "content": "hi"}],
api_key="sk-invalid",
num_retries=0,
)
except routehub.AuthenticationError as error:
print(type(error).__name__, error.status_code, error.llm_provider)Run in production
No global state
Retries, TLS verification, timeouts and parameter dropping are arguments to each call. One tenant's settings never leak into another's, which matters when many agents or teams share a process.
- drop_params, num_retries, ssl_verify, request_timeout and set_verbose on every call
- Pass your own OpenAI or httpx client for a proxy or mTLS
- Clients are pooled per configuration, so per-call settings cost one hash lookup
- set_verbose never prints API keys or message content
import routehub
# Agent A: behind a corporate proxy with a private CA bundle
routehub.completion(
model="gpt-5.4-mini",
messages=messages,
ssl_verify="/etc/ssl/corp-ca.pem",
num_retries=5,
request_timeout=30,
)
# Agent B, same process: defaults, untouched by Agent A
routehub.completion(model="claude-sonnet-4-6", messages=messages)import litellm
# Module globals: every caller in the process now
# shares these settings, including other tenants
litellm.ssl_verify = False
litellm.num_retries = 5
litellm.drop_params = True
litellm.completion(model="gpt-5.4-mini", messages=messages)Run in production
Model catalog and costs
Context windows, output limits, per-token prices and capabilities from OpenRouter's live model list, plus every model each provider serves from its own API. Both cache for five minutes, with no bundled data file to go stale.
- get_model_info returns per-token prices, so spend can be computed per call
- Usage comes back in one shape for every provider, including cached and reasoning tokens
- register_model adds private, fine-tuned or self-hosted models
- get_all_models queries every configured provider concurrently
import routehub
info = routehub.get_model_info("claude-sonnet-4-6")
print(info["max_input_tokens"], info["max_output_tokens"])
response = routehub.completion(
model="claude-sonnet-4-6", messages=messages
)
usage = response.usage
cost = (
usage.prompt_tokens * info["input_cost_per_token"]
+ usage.completion_tokens * info["output_cost_per_token"]
)
print(f"${cost:.6f}")from routehub.get_all_models import get_all_models, get_model
# Every provider with a key configured, plus OpenRouter
models = get_all_models()
chat = [m for m in models if m["type"] == "chat"]
print(len(chat), "chat models")
get_all_models(["anthropic", "gemini"]) # selected providers
get_model("anthropic", "claude-sonnet-4-6") # one modelimport routehub
routehub.register_model({
"acme-support-ft": {
"max_input_tokens": 32768,
"max_output_tokens": 4096,
"supports_function_calling": True,
"input_cost_per_token": 0.000002,
}
})
print(routehub.supports_function_calling("acme-support-ft"))Run in production
Testing without a provider
mock_response returns a real ChatCompletion, or a stream when stream=True, with no network call and no API key. Pass an exception to rehearse failure handling.
- Runs in CI with no secrets configured
- Exceptions rehearse outages, timeouts and retry paths
- RouteHub itself ships with 231 offline tests
import routehub
response = routehub.completion(
model="gpt-5.4-mini",
messages=[{"role": "user", "content": "Is the deploy approved?"}],
mock_response="Approved.",
)
print(response.choices[0].message.content)
print(type(response).__name__)import routehub
try:
routehub.completion(
model="gpt-5.4-mini",
messages=[{"role": "user", "content": "Is the deploy approved?"}],
mock_response=TimeoutError("simulated outage"),
)
except TimeoutError as error:
print(type(error).__name__, error)Migrate
Drop-in for LiteLLM
RouteHub keeps LiteLLM's function names and module paths, so moving over is mostly a change of import. Settings move from module globals to arguments.
- routehub.utils and routehub.exceptions mirror litellm's modules
- Responses are OpenAI SDK objects: use attributes or .model_dump() instead of dictionary keys
- max_tokens is sent to OpenAI as max_completion_tokens, which reasoning models require
- litellm-only arguments such as metadata and caching are accepted and ignored
import litellm
from litellm import completion
from litellm.exceptions import RateLimitError
litellm.drop_params = True
litellm.num_retries = 3
response = completion(model="claude-sonnet-4-6", messages=messages)
print(response["choices"][0]["message"]["content"])from routehub import completion
from routehub.exceptions import RateLimitError
response = completion(
model="claude-sonnet-4-6",
messages=messages,
drop_params=True,
num_retries=3,
)
print(response.choices[0].message.content)Get started
Start a new project, move off LiteLLM, or teach your coding agent
RouteHub is free and open source under Apache 2.0. Pick the path that matches where you are today.
New project
Install the package, export a key for any provider, and make your first call.
pip install routehubMigrating from LiteLLM
Same function names and module paths. Swap the import, then move global settings into each call.
- from litellm import completion+ from routehub import completion- litellm.num_retries = 3- completion(model=model, messages=messages)+ completion(model=model, messages=messages, num_retries=3)
Coding agents
RouteHub ships an Agent Skill that shows Claude Code and other coding agents how to call it correctly, including which LiteLLM habits break.
git clone https://github.com/The-Swarm-Corporation/RouteHub.git mkdir -p ~/.claude/skills cp -r RouteHub/skills/routehub ~/.claude/skills/
Paper
The design and every benchmark, written up in full
FAQ
Questions about RouteHub
What Python developers ask before they switch LLM gateways.
What is RouteHub?
RouteHub is an open-source LLM gateway for Python from the Swarms team. One completion() function calls OpenAI, Anthropic, Gemini, Groq, xAI, DeepSeek, OpenRouter, Mistral, Together, Fireworks, Azure OpenAI, Ollama, vLLM and any OpenAI-compatible server, and every response comes back as the official OpenAI SDK's ChatCompletion type. Install it with pip install routehub.
Read the launch postIs RouteHub a drop-in replacement for LiteLLM?
For most Python code, yes. RouteHub keeps LiteLLM's function names and module paths, including completion, acompletion, embedding and routehub.utils.get_model_info, so the main change is the import. Settings that LiteLLM keeps in module globals, such as retries and timeouts, become arguments to each call.
Migrate from LiteLLM step by stepHow much faster is RouteHub than LiteLLM?
In the RouteHub paper's benchmarks against LiteLLM 1.104.0, RouteHub imports in 3.2 ms against 1,235 ms (383x faster), reaches a first response from a fresh process in 267 ms against 1,464 ms (5.5x faster), peaks at 54 MiB of memory against 211 MiB, and adds 9 µs to each OpenAI call against 647 µs. LiteLLM's aiohttp transport is ahead on async throughput at 128 requests in flight.
Why is LiteLLM slow?Which LLM providers does RouteHub support?
More than 20, including OpenAI, Anthropic, Google Gemini, Groq, xAI, DeepSeek, OpenRouter, Mistral, Together, Fireworks, Azure OpenAI, Ollama and vLLM. Any server that speaks the OpenAI API works through api_base, including private endpoints and local servers. Native Bedrock and Vertex AI adapters are on the roadmap.
Does RouteHub support streaming, async, tool calling and structured output?
Yes. stream=True yields ChatCompletionChunk objects, acompletion takes the same arguments as completion for concurrent requests, tools defined once in the OpenAI format work on every provider, and passing a pydantic model returns JSON that validates against it. For Claude, RouteHub translates tool definitions, calls and results to Anthropic's format and back.
Does RouteHub support Claude prompt caching and extended thinking?
Yes. Anthropic's OpenAI-compatible endpoint drops prompt caching, thinking output and cached-token usage, so RouteHub calls Anthropic's native Messages API instead and keeps all three. In the paper's benchmarks, RouteHub's Claude path was also faster than the official Anthropic SDK for both regular and streaming calls.
Use Claude in OpenAI format from PythonDoes RouteHub include a proxy server like LiteLLM's?
No. RouteHub is a Python library that runs inside your process. It has no proxy server, budgets, spend logging, callbacks or router. Teams that depend on those features can keep the LiteLLM proxy or choose a self-hosted gateway, and our comparison of LiteLLM alternatives covers both options.
Compare LiteLLM alternativesIs RouteHub free?
Yes. RouteHub is free and open source under the Apache 2.0 license and runs on Python 3.10 and later. It has three direct dependencies (openai, pydantic and tiktoken) and installs 21 packages in total, against 58 for LiteLLM. You pay only your model providers for the tokens you use.
Compare Python LLM gateways