Streaming
stream true on every metered endpoint, how settlement works on the terminal chunk, and what to do when a stream breaks.
Every metered endpoint streams. Set stream: true and you get server-sent events —
/v1/chat/completions and /v1/completions in the OpenAI chunk shape, so a stock OpenAI
client iterates them unchanged; /v1/messages and /v1/responses in their own native
Anthropic and OpenAI Responses event shapes, respectively.
stream = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Write a haiku about prepaid inference."}],
max_tokens=100,
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)Billing on a stream
The prepaid flow is the same as a non-streamed call, and the difference is only when the books close:
- The worst-case hold is reserved before the first byte is sent.
- Chunks stream to you.
- The terminal chunk carries the final
usage, and settlement happens there.
So a 402 for insufficient balance arrives before the stream starts, never halfway
through one. If a stream begins, the hold cleared.
Read usage off the terminal chunk
Intermediate chunks do not carry a settled usage. Take token counts and
usage.cost from the last chunk of the stream. Reading them earlier gets you a
partial figure or nothing at all.
When a stream breaks
If the connection drops mid-stream, you have a partial completion and ModelBeat has an unsettled hold. The hold is reconciled on our side rather than being stranded against your balance — you are charged for what was actually generated, not for the reservation.
There is no resume. Reissue the request.
Errors during a stream
An error raised before the stream starts uses the ordinary envelopes and status codes in Errors — a normal JSON body, not an SSE frame. That covers authentication, rate limits, balance, and routing refusals, all of which are decided up front.
Every response carries X-Request-Id, streamed or not. Log it. See
Errors.
Routing
What auto does, the intelligence tiers, the gates a request has to clear, what happens on a fallback, and why model identity is not on the response.
Capabilities
What actually works today beyond the base request shape — tool calling, vision, and the in-repo SDKs — each confirmed by live testing, not assumed.