Limits
Current constraints. Four metered endpoints, model auto only, per-key rate limits, and the request shapes we do not support.
Read this before you design an integration
The constraints below are deliberate, and they are the kind of thing that costs an afternoon if you find them late. ModelBeat is in private beta: this surface can change, and it will change by growing rather than by breaking what is documented here.
Four metered endpoints
| Endpoint | Status |
|---|---|
POST /v1/chat/completions | Supported (streaming and non-streaming) |
POST /v1/completions | Supported (streaming and non-streaming) |
POST /v1/messages (Anthropic shape) | Supported (streaming and non-streaming), as of 2026-09-01. Accepts Anthropic's wire format — ModelBeat's own routing still decides which model serves the request. Prompt-cache tokens are billed at the regular input rate, not Anthropic's cheaper cache-read rate: you may see a slightly higher charge than calling Anthropic directly on a cache hit. |
POST /v1/responses (OpenAI Responses shape) | Supported (streaming and non-streaming), as of 2026-09-01. |
That is the entire metered surface. GET /v1/models is supported but unmetered. The
rest of the gateway's paths are not part of the API, and since 2026-07-28 they fail
loudly rather than appearing to work:
| Path | What you get |
|---|---|
/anthropic/v1/messages | 501. A separate, still-unsupported mount — use /v1/messages instead. |
/v1/responses/input_tokens, /v1/responses/compact, /v1/messages/{path} (batches) | 501. Sub-resources are not part of the supported surface. |
/v1/embeddings | Supported and functional — a provider is called and returns real embeddings — but not metered: no balance hold, no usage.cost, no usage-history entry. Unlike every other endpoint, model takes a real provider model id on input, not auto/a tier — there is no tier abstraction for embeddings. Unlike a metered response, extra_fields.routing_info/provider in this endpoint's body carry the real provider/model — that redaction only runs on the metered/settled path, which this endpoint never enters. See the reference for the exact caveats before you build on it. |
/v1/models | Supported, authenticated, rate-limited, never metered. During beta it lists the three intelligence tiers and never a model id. See the reference. |
/v1/feedback | Supported, authenticated and rate-limited, never metered. Identify the request being scored by its X-Request-Id header, not the response body's id field. See the reference. |
/v1/balance | Supported, authenticated, rate-limited, never metered. A standalone, point-in-time balance breakdown. See the reference. |
/v1/usage | Supported, authenticated, rate-limited, never metered. Account-level aggregate spend/request/token count over a date window (defaults to the trailing 30 days), merged with a balance snapshot. See the reference. |
/v1/requests | Supported, authenticated, rate-limited, never metered. Paginated per-request usage history, same date-window default as /v1/usage. See the reference. |
/v1/requests/{id} | Supported, authenticated, rate-limited, never metered. Single request lookup by X-Request-Id — same identifier convention as /v1/feedback, 404 for a request that never existed or belongs to another tenant. See the reference. |
/v1/async/* | 404. Not on the gateway's product surface at all: the edge gate is default-deny, and these paths are not in the allow-set. |
every /api/* path | 404. This is the vendored gateway's own admin surface, not part of ModelBeat. |
Anything not in the API reference is not part of the API. It may be removed, gated, or changed without notice, and calls to it are not covered by any correctness or billing guarantee.
model does not take a model id
model takes auto or an intelligence tier
(modelbeat-advanced, modelbeat-standard, modelbeat-fast). A tier names a capability
band, not a model. You cannot name, pin, prefer, or exclude a specific model, and the
only catalogue GET /v1/models lists is those three tiers. During the current private beta,
which model served a request is not exposed in the response body — model returns the
intelligence tier, not a model id, and routing_info is rewritten to that tier-only shape
before settlement (ADR-0067). Provider-level response headers on some routes are a separate,
known gap tracked outside this promise as RA-010. See Routing for why the
response-body redaction is a beta constraint rather than a permanent one, and for how to get
a per-request record when you need one.
Rate limits
Rate limiting is per API key, measured in requests per minute, and applies to every
/v1 path — including the unmetered ones, because an unmetered request is still
inference somebody pays for.
Over the limit you get 429:
{ "error": "modelbeat: rate limit exceeded for this api key" }with a Retry-After header carrying the number of seconds until the window resets.
Honour it; retrying sooner just burns another refusal.
A 429 is refused before routing and before the balance hold, so nothing is charged
and no reservation is left stranded. It does not appear in your usage history, because no
model was called.
Every /v1 response from a rate-limited key — success or 429 — also carries:
| Header | Meaning |
|---|---|
X-RateLimit-Limit | Requests per minute allowed for this key. |
X-RateLimit-Remaining | Requests left in the current window. 0 on a 429. |
X-RateLimit-Reset | Seconds until the window resets — not an epoch timestamp. On a 429 this always matches Retry-After. |
Use X-RateLimit-Remaining to pace yourself before you get refused, not just after.
Beta default: no per-key limit
Keys issued to date carry no rate limit, so in practice you will not see a 429
today. That is a beta convenience, not a guarantee. Handle 429 and Retry-After now
rather than discovering the limit the day it is switched on.
Rate limits are separate from your prepaid balance. A 429 says "too fast"; a 402 says
"not enough balance". See Costs.
Other current constraints
response_format/ structured outputs is not supported. Bothjson_schemaandjson_objectare silently dropped today — the request succeeds, but you get free-text markdown back, not the schema-conformant output you asked for, and nothing in the response signals that it happened. Do not build on this until it is confirmed fixed.- One key class. Only
mb_live_keys exist. There is no separate admin or management key, and no public management API. Key, balance, and usage administration happens in the console. usagemay benull. When the serving provider reports no usage, the field is present and null rather than omitted. Guard it.- Undocumented response members are not contract. Responses carry
model(the serving tier) and a ModelBeatextra_fieldsobject, of which only the members the reference documents are contract. Ignore the rest: they are unstable, may be redacted or removed without notice, and carry no correctness guarantee. - No published SLA. Private beta. See Support.