Streaming

Streaming

Set stream: true for server-sent events in exactly OpenAI's format. Chunks arrive as data: lines and the stream ends with data: [DONE].

python
stream = client.chat.completions.create(
    model="cbcn/glm-5.3",
    messages=[{"role": "user", "content": "Count to five"}],
    stream=True,
)

for chunk in stream:
    delta = chunk.choices[0].delta.content if chunk.choices else None
    if delta:
        print(delta, end="", flush=True)

Usage on the final frame

We always request usage reporting from the upstream model, so the last frame before [DONE] carries a usage object with the real token counts. You do not need to set stream_options yourself. If you do, it is respected.

text
data: {"choices":[{"delta":{"content":"1"}}],"object":"chat.completion.chunk"}

data: {"choices":[],"usage":{"prompt_tokens":9,"completion_tokens":25,"total_tokens":34}}

data: [DONE]

Cancelling a stream

Disconnecting stops generation. You are billed for the tokens produced up to that point, which the upstream model has already generated and charged us for. Requests that fail before producing any tokens are never billed.

If you sit behind a reverse proxy, disable response buffering for this endpoint or chunks will be held back and appear to arrive all at once. We send X-Accel-Buffering: no to help nginx do the right thing.