Guide

SillyTavern with AtmoRouter: Frontend Characters, Frontier Backends

Connect SillyTavern to any AtmoRouter model through its OpenAI-compatible Chat API. Full setup, presets for streaming and context size, and how to keep character-card costs predictable.

Oct 6, 20265 min readAtmoRouter Team

SillyTavern is where a lot of people actually live with LLMs — character cards, long chats, world-building — and its backend is pluggable by design. Pointing it at AtmoRouter gives every character the same catalogue everyone else uses: one key, per-model rates printed up front, and the same OpenAI-compatible Chat Completions API your other tools already speak.

Setup

  1. 1.In SillyTavern, open API Connections (top-right plug icon).
  2. 2.API: Chat Completion. Chat Completion Source: Custom (OpenAI-compatible).
  3. 3.Custom Endpoint: https://atmorouter.dev/v1 — no /chat/completions suffix; SillyTavern appends it.
  4. 4.API Key: your AtmoRouter key from the dashboard.
  5. 5.Load the model list, and pick any AtmoRouter model id (e.g. zai/glm-5.3).

That is the whole integration. No proxy, no local relay.

Presets that matter

  • Streaming: leave it on. AtmoRouter streams server-sent events on every model, and character chats read very differently at 40 tokens vs 4,000.
  • Context size: SillyTavern will happily stuff 100k+ tokens of chat history into a request. Every turn re-sends the context, so a long chat costs like a long document each time. Two defenses:

- Cap context to what the card actually needs (8-16k is plenty for most characters). - Pick models where cached input bills at 10% of the input rate — a returning context is mostly cache hits, and the bill follows.

  • NSFW/uncensored note: filtering is a property of each upstream model, not of the gateway. Check the model's page for family notes before writing a card around it.

Cost sanity for character chats

A 10k-token context, 300-token replies, 100 turns a day: roughly $0.004-0.01 per turn on mid-tier models with caching — a few cents to about a dollar a day for heavy use. The rate card is on every model page, so this arithmetic takes thirty seconds before you commit a card to a model.

Why route this through one gateway at all

Because the backend you want at 2 a.m. for a creative chat is not always the backend you want at 9 a.m. for code review. Same key, same endpoint, different model string — SillyTavern, your IDE, your scripts all draw from one balance, and the per-model price is printed where you pick it.

Mentioned in this post

Start building today

Create an account, grab a key, and make your first call in minutes.