Features

Built to be boring in production

An inference gateway earns trust by being predictable: exact billing, honest errors, and no surprises when a model goes down. Here is what that means in practice.

One key, every model

A single ar- key reaches the whole catalogue. No per-provider signups, no separate billing relationships, no juggling five dashboards.

  • Drop-in OpenAI replacement
  • Anthropic Messages API supported
  • Unlimited keys per account

Routed to healthy capacity

Requests land on capacity that is up and within budget. Exhausted or degraded sources are skipped automatically, so you are not the one discovering an outage.

  • Automatic failover across capacity
  • Per-family health tracked continuously
  • No manual re-rolls

Exact per-token metering

Billing is computed from the upstream usage report, never estimated from characters. The number you are charged is the number you consumed.

  • Input and output priced separately
  • Usage reported on the final stream frame
  • X-Atmorouter-Cost on every response

Streaming that streams

Full SSE passthrough in both OpenAI and Anthropic formats, with no buffering layer in between. Stop the stream and you stop paying.

  • Byte-for-byte SSE passthrough
  • Anthropic event protocol supported
  • Cancel safely mid-generation

Keys with guardrails

Issue a key per application, cap its rate, and restrict which models it may reach. Revoking one never disturbs the others.

  • Per-key rate limits
  • Per-key model allowlists
  • Instant revocation

Usage you can audit

Every request is logged with tokens, latency, cost and what the same call would have cost at official list prices.

  • Per-model and per-family breakdown
  • Latency percentiles
  • Saved-versus-official on every window

Measured, not claimed

These come straight from our own request log. A failed upstream request is never billed to you.

—
Success rate, 7d
—
Median latency
0
Requests served
0
Tokens routed

Transparent billing

Prepaid balance, per-token deduction, and a receipt on every single response.

What you get

  • One key for OpenAI- and Anthropic-compatible endpoints
  • Every models across all families
  • Real-time balance deduction per request
  • Per-key rate limits and model restrictions
  • Usage dashboard with cost and latency breakdown

How billing works

  • →Deposit USDC on Solana, or top up with QRIS
  • →Input and output tokens counted separately
  • →Cached input bills at 10% of the input rate
  • →cost = ((in − cached) × input_price + cached × input_price × 0.10 + out × output_price) / 1M
  • →Balance deducted the moment a request settles
  • →A failed upstream request is never billed
  • →Empty balance blocks new calls, never overdraws

Start building today

Create an account, grab a key, and make your first call in minutes.