Logo RouteroAI
How it works

One API call,
four decisions.

Every request that hits Routero AI passes through a deterministic pipeline. Policy first, then routing, then accounting, then audit — sub-50ms added latency, all of it explainable.

01

Receive the request.

Your app calls api.routero.ai/v1/chat/completions using the OpenAI SDK you already have. We accept any OpenAI-compatible payload — streaming, tools, vision, JSON mode. Beyond chat, embeddings, rerank, image generation, and audio endpoints are exposed under the same API.

app.py
from openai import OpenAI

client = OpenAI(
    base_url="https://api.routero.ai/v1",
    api_key="rt_live_...",
)

resp = client.chat.completions.create(
    model="customer-support",    # or "gpt-5.4", "claude-sonnet-4.6", ...
    messages=[...],
)
02

Run the policy gate.

Four checks fire in parallel: identity (RBAC, workspace scope), content (prompt-injection score, PII detection), model (is the requested model allowed for this caller), and budget (does this request still fit). Any failure short-circuits with a structured error.

Inbound
openai.chat()
→
Identity · team-finance
Content · 0.02 injection score
Model · customer-support route allowed
Budget · $8,412 / $22,000
→
Decision
ROUTE
03

Pick the provider.

Routero AI scores eligible providers by health, latency, price, residency, and recent error rate. The chosen provider is called with full streaming pass-through. If it fails or rate-limits, the next eligible provider is tried automatically.

★
Anthropic claude-sonnet-4.6
P50 142ms · $0.003/1k · OK
SELECTED
2
OpenAI gpt-5.4
P50 178ms · $0.005/1k · 429s
fallback
3
Gemini 2.0-flash
P50 121ms · $0.002/1k · region-mismatch
ineligible
4
Mistral large-2
P50 198ms · $0.004/1k · OK
fallback
04

Account & audit.

Token counts and dollar cost are debited from the budget atomically. The full decision — every check, the chosen provider, the alternatives considered, the response shape — is written to a centralized audit log in real time, queryable and exportable at any moment.

Budget update
team-finance / aug+$0.0124
remaining$13,587.84
Audit entry
{
  "req": "req_8b2f...",
  "user": "u_pa9",
  "route": "customer-support",
  "chose": "anthropic:claude-sonnet-4.6",
  "checks": ["id","content","model","budget"],
  "latency_ms": 142,
  "tokens": { "in": 412, "out": 198 }
}

The Four Building Blocks.

Everything you configure in Routero AI is one of these. Compose them to model your real organization.

Three deployment modes.

Same control plane, different trust boundaries — pick the one your security team is comfortable with.

Routero AI Cloud

Multi-tenant, multi-region. SOC 2 controls (in progress). Start in 60 seconds with an API key.

  • · Managed infrastructure
  • · Auto-scaling
Coming soon

Self-hosted (VPC)

Run Routero AI inside your VPC. Zero third-party data paths. Docker Compose and containerized delivery.

  • · In-VPC deployment
  • · Air-gapped option
  • · Data stays inside your network

See it on your stack.

A 30-minute walkthrough with a solutions engineer — bring your provider list and we'll map it live.