Features
An inference gateway earns trust by being predictable: exact billing, honest errors, and no surprises when a model goes down. Here is what that means in practice.
A single ar- key reaches the whole catalogue. No per-provider signups, no separate billing relationships, no juggling five dashboards.
Requests land on capacity that is up and within budget. Exhausted or degraded sources are skipped automatically, so you are not the one discovering an outage.
Billing is computed from the upstream usage report, never estimated from characters. The number you are charged is the number you consumed.
Full SSE passthrough in both OpenAI and Anthropic formats, with no buffering layer in between. Stop the stream and you stop paying.
Issue a key per application, cap its rate, and restrict which models it may reach. Revoking one never disturbs the others.
Every request is logged with tokens, latency, cost and what the same call would have cost at official list prices.
These come straight from our own request log. A failed upstream request is never billed to you.
Prepaid balance, per-token deduction, and a receipt on every single response.
Create an account, grab a key, and make your first call in minutes.