Routero AI is an enterprise control plane that sits between your application and every AI provider. One OpenAI-compatible API gets you access to 100+ models — with built-in failover, governance, spend controls, and audit trails.
Think of it as the "Stripe Radar + Datadog + IAM" layer for your AI traffic.
If your app already uses the OpenAI SDK, integration is one line — change base_url to https://api.routero.ai/v1. Most teams have their first routed request flowing in under 10 minutes.
No. Routero AI passes your prompts through verbatim. You can call a custom route alias (e.g. customer-support) and let routing strategies pick the model, or call gpt-5.4 or claude-sonnet-4.6 directly.
OpenAI, Anthropic, Google Gemini, Azure OpenAI, AWS Bedrock, Mistral, Cohere, Groq, Together, Fireworks, Replicate, Hugging Face Inference, and any self-hosted OpenAI-compatible endpoint (vLLM, Ollama, TGI). Full list: 100+ models across 20+ provider clouds.
Yes. The Free plan includes 1M routed tokens per month, no card required. You can use Routero AI's API keys (pay provider pricing 1:1, no markup) or bring your own.
Every route has an ordered list of provider+model candidates. If the primary fails (timeout, 5xx, rate limit, or health-probe degraded), Routero AI retries the next candidate transparently. Your application receives a single successful response. Streaming requests that fail before output starts switch over automatically too.
Routes can select models with cost-first, latency-first, or load-balancing strategies — scoring candidates on health, recent error rates, and cooldown state. Semantic model selection is also supported. You define your own route aliases (e.g. customer-support) and attach strategies to them.
Yes. Define a route with a traffic split (e.g. 70% GPT-5.4, 30% Claude). Routero AI stamps the chosen provider on every request and records it in your audit log — feed that into your experimentation tooling.
Yes — Server-Sent Events stream straight through with no buffering. Tool calls, vision, JSON mode, and structured outputs all pass through unmodified.
The same API also exposes embeddings, rerank, image generation, audio, files, and batch endpoints. The gateway supports MCP (Model Context Protocol) and A2A agent endpoints, so agent traffic is governed on the same control plane as model traffic. Response caching and semantic caching are supported — cache hits cut token costs significantly.
< 50ms P99 for the routing decision itself. In practice, P50 added latency is around 8–12ms. Routero AI is colocated with provider POPs, so geographic overhead is minimal.
Regional SaaS nodes keep requests and logs in-region and only peer with in-region providers; self-hosted models are only called from inside your VPC. Combined with PII detection and masking, this supports your compliance review. Note: this page is not legal advice — final compliance conclusions should come from your legal team.
SOC 2 Type II is in progress (WIP) — audit underway, with the report expected to be available under NDA once complete. HIPAA-eligible plans available on Enterprise.
No. Routero AI never trains on, retains, or shares your prompts or completions. Request payloads are processed in memory and discarded; only metadata (token counts, latency, decisions) is persisted for audit and billing.
OIDC single sign-on is available from the Growth plan — compatible with Microsoft Entra, Google Workspace, and any OIDC IdP (Okta, Auth0, Keycloak, custom). SCIM user provisioning is coming soon.
Every routing decision: caller identity, route requested, policy checks fired, provider chosen, alternatives considered, latency, token usage, and a request-id correlation hash. Exportable to Datadog or Azure Sentinel, or via OpenTelemetry into ELK, Loki, and your existing log platform.
Self-hosted (VPC) deployment runs in your own VPC or data center — Docker Compose and containerized delivery, with an air-gapped option for regulated workloads. Enterprise customers can join the early-access program today. In the meantime, Single-tenant Cloud (dedicated cluster in your region, managed by us) is available now.
You have two options, and you can mix them per workspace:
1. Use Routero AI's keys. We provision API access to all supported providers. You pay provider-list prices, passed through 1:1, on a single consolidated Routero AI invoice. We never mark up provider costs for profit — every charge is accountable and tied to an operational purpose.
2. Bring your own keys (BYOK). Plug in your existing provider keys. Your provider invoices stay exactly as they are today; you only pay Routero AI for the control plane subscription.
Most teams start with our keys for speed and move portions to BYOK once they've negotiated direct enterprise pricing with specific providers.
No. When you use our keys, token costs are passed through at provider list price — penny-for-penny. Our revenue comes entirely from the control-plane subscription (routing, governance, audit, spend dashboards). Every line item on your bill is operational and verifiable against the provider's published rates.
One inbound API call that Routero AI evaluates and routes. Retries to fallback providers don't count as additional requests. Streaming counts as one request regardless of token count.
Yes — 20% off Growth and Enterprise on annual contracts. Multi-year commitments available.
Tag every request with a cost-center label (team, project, customer-id). The chargeback dashboard rolls token spend and provider cost into per-label reports, exportable as CSV.
US (us-east, us-west), EU (eu-west, eu-central), and APAC (ap-southeast). Enterprise customers can pin workloads to a specific region.
Yes. Any OpenAI-compatible endpoint (vLLM, Ollama, TGI, custom) can be added as a provider. Route to it like any first-party model.
SaaS is deployed across multiple availability zones by default; the gateway is stateless and scales horizontally, and provider health checks remove failing nodes automatically. Your apps keep using the standard OpenAI SDK — and for extreme scenarios you can keep a provider-native direct config as an escape hatch for critical traffic.
Try a different keyword, or ask our team.
Book a 30-minute call with a solutions engineer. We answer technical questions live.
Book a demo