API reference
11 endpoints, generated from the OpenAPI contract.
ModelBeat is OpenAI-compatible: point an existing SDK at https://api.modelbeat.ai/v1 and change the key. Streaming is supported, and model takes auto or an intelligence tier. See Routing and Limits.
Create a chat completionOpenAI-compatible chat completion.
ModelBeat reserves a worst-case cost hold before routing, then settles it
against realized usage after the provider responds. A `402` means the hold
exceeded the available prepaid balance.
Create a text completionOpenAI-compatible text completion.
Create a message (Anthropic shape)Anthropic Messages API-compatible endpoint. Accepts Anthropic's wire format;
ModelBeat's own catalog-driven routing still decides which model actually serves the
request — sending this shape does not pin you to an Anthropic-served model, and does
not require one to exist in the catalog. Prompt-cache tokens
(`cache_creation_input_tokens`/`cache_read_input_tokens`) are billed at the regular
input rate, not Anthropic's cheaper cache-read rate.
Create a response (OpenAI Responses shape)OpenAI Responses API-compatible endpoint. Same routing/billing semantics as
`/chat/completions` and `/messages` — a different wire format for the same
catalog-driven model pool.
List the intelligence tiersLists the values `model` accepts, in the OpenAI `/v1/models` list shape, so a stock
OpenAI SDK's `models.list()` works unchanged. Authenticated and rate-limited like
every `/v1` path, but never metered — no balance hold is taken.
During the private beta this is a tier list, not a model catalogue: it returns
`modelbeat-advanced`, `modelbeat-standard` and `modelbeat-fast`, and never a model
id. Tier names are ModelBeat's own abstraction over the pool, so listing them
discloses no model identity (ADR-0070).
Create embeddingsOpenAI-compatible embeddings. Real inference — a provider is called and paid for
this request — but this path is **not metered by ModelBeat**: no balance hold is
taken, no `usage.cost` is attached, and the call does not appear in usage
history. Do not build a production integration that depends on this endpoint's
cost or usage being tracked.
Get the prepaid balanceA standalone, point-in-time balance check — the full breakdown (held, purchased,
bonus, signup offer), for a "do I have enough money to fire this batch right now"
check before a caller submits a large amount of work. Authenticated and
rate-limited like every `/v1` path, but never metered — no balance hold is taken.
Get an aggregate spend/usage summaryAccount-level aggregate — total spend, request count, and token count over a date
window — merged server-side with a balance snapshot. Authenticated and
rate-limited like every `/v1` path, but never metered.
Omit both `from` and `to` and the window defaults to the trailing 30 days; the
response always echoes back the window actually queried, defaulted or not. For a
per-request list instead of an aggregate, see `GET /requests`.
List past requestsPaginated, per-request usage history — an OpenAI-style list envelope. Authenticated
and rate-limited like every `/v1` path, but never metered.
Omit both `from` and `to` and the window defaults to the trailing 30 days, echoed
back in the response the same way `GET /usage` does. For a single request by id,
see `GET /requests/{id}`.
Get a single past requestLook up one request by id. Authenticated and rate-limited like every `/v1` path,
but never metered.
**`id` must be the `X-Request-Id` response header value from the original
request, not the response body's own `id` field.** Same identifier convention as
`POST /feedback`.
Submit an outcome signal for a prior requestExplicit outcome signal (`"up"` or `"down"`) for a request you made earlier,
keyed by that request's id.
**`request_id` must be the `X-Request-Id` response header value from the original
request, not the response body's own `id` field.** The two are different
identifiers, and only the header value is recognized here — the body's `id` returns
`404`.
Authenticated and rate-limited like every `/v1` path, but never metered: no balance
hold, no usage row, no routing.
Everything else the gateway exposes is not part of the API.