Streaming
Streaming
Set stream: true for server-sent events in exactly OpenAI's format. Chunks arrive as data: lines and the stream ends with data: [DONE].
python
stream = client.chat.completions.create(
model="cbcn/glm-5.3",
messages=[{"role": "user", "content": "Count to five"}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content if chunk.choices else None
if delta:
print(delta, end="", flush=True)Usage on the final frame
We always request usage reporting from the upstream model, so the last frame before [DONE] carries a usage object with the real token counts. You do not need to set stream_options yourself. If you do, it is respected.
text
data: {"choices":[{"delta":{"content":"1"}}],"object":"chat.completion.chunk"}
data: {"choices":[],"usage":{"prompt_tokens":9,"completion_tokens":25,"total_tokens":34}}
data: [DONE]Cancelling a stream
Disconnecting stops generation. You are billed for the tokens produced up to that point, which the upstream model has already generated and charged us for. Requests that fail before producing any tokens are never billed.
If you sit behind a reverse proxy, disable response buffering for this endpoint or chunks will be held back and appear to arrive all at once. We send
X-Accel-Buffering: no to help nginx do the right thing.