The surface is narrow on purpose: two metered endpoints and model: "auto". Model identity is not exposed.
ModelBeatbeta

Limits

Current constraints. Four metered endpoints, model auto only, per-key rate limits, and the request shapes we do not support.

Read this before you design an integration

The constraints below are deliberate, and they are the kind of thing that costs an afternoon if you find them late. ModelBeat is in private beta: this surface can change, and it will change by growing rather than by breaking what is documented here.

Four metered endpoints

EndpointStatus
POST /v1/chat/completionsSupported (streaming and non-streaming)
POST /v1/completionsSupported (streaming and non-streaming)
POST /v1/messages (Anthropic shape)Supported (streaming and non-streaming), as of 2026-09-01. Accepts Anthropic's wire format — ModelBeat's own routing still decides which model serves the request. Prompt-cache tokens are billed at the regular input rate, not Anthropic's cheaper cache-read rate: you may see a slightly higher charge than calling Anthropic directly on a cache hit.
POST /v1/responses (OpenAI Responses shape)Supported (streaming and non-streaming), as of 2026-09-01.

That is the entire metered surface. GET /v1/models is supported but unmetered. The rest of the gateway's paths are not part of the API, and since 2026-07-28 they fail loudly rather than appearing to work:

PathWhat you get
/anthropic/v1/messages501. A separate, still-unsupported mount — use /v1/messages instead.
/v1/responses/input_tokens, /v1/responses/compact, /v1/messages/{path} (batches)501. Sub-resources are not part of the supported surface.
/v1/embeddingsSupported and functional — a provider is called and returns real embeddings — but not metered: no balance hold, no usage.cost, no usage-history entry. Unlike every other endpoint, model takes a real provider model id on input, not auto/a tier — there is no tier abstraction for embeddings. Unlike a metered response, extra_fields.routing_info/provider in this endpoint's body carry the real provider/model — that redaction only runs on the metered/settled path, which this endpoint never enters. See the reference for the exact caveats before you build on it.
/v1/modelsSupported, authenticated, rate-limited, never metered. During beta it lists the three intelligence tiers and never a model id. See the reference.
/v1/feedbackSupported, authenticated and rate-limited, never metered. Identify the request being scored by its X-Request-Id header, not the response body's id field. See the reference.
/v1/balanceSupported, authenticated, rate-limited, never metered. A standalone, point-in-time balance breakdown. See the reference.
/v1/usageSupported, authenticated, rate-limited, never metered. Account-level aggregate spend/request/token count over a date window (defaults to the trailing 30 days), merged with a balance snapshot. See the reference.
/v1/requestsSupported, authenticated, rate-limited, never metered. Paginated per-request usage history, same date-window default as /v1/usage. See the reference.
/v1/requests/{id}Supported, authenticated, rate-limited, never metered. Single request lookup by X-Request-Id — same identifier convention as /v1/feedback, 404 for a request that never existed or belongs to another tenant. See the reference.
/v1/async/*404. Not on the gateway's product surface at all: the edge gate is default-deny, and these paths are not in the allow-set.
every /api/* path404. This is the vendored gateway's own admin surface, not part of ModelBeat.

Anything not in the API reference is not part of the API. It may be removed, gated, or changed without notice, and calls to it are not covered by any correctness or billing guarantee.

model does not take a model id

model takes auto or an intelligence tier (modelbeat-advanced, modelbeat-standard, modelbeat-fast). A tier names a capability band, not a model. You cannot name, pin, prefer, or exclude a specific model, and the only catalogue GET /v1/models lists is those three tiers. During the current private beta, which model served a request is not exposed in the response body — model returns the intelligence tier, not a model id, and routing_info is rewritten to that tier-only shape before settlement (ADR-0067). Provider-level response headers on some routes are a separate, known gap tracked outside this promise as RA-010. See Routing for why the response-body redaction is a beta constraint rather than a permanent one, and for how to get a per-request record when you need one.

Rate limits

Rate limiting is per API key, measured in requests per minute, and applies to every /v1 path — including the unmetered ones, because an unmetered request is still inference somebody pays for.

Over the limit you get 429:

{ "error": "modelbeat: rate limit exceeded for this api key" }

with a Retry-After header carrying the number of seconds until the window resets. Honour it; retrying sooner just burns another refusal.

A 429 is refused before routing and before the balance hold, so nothing is charged and no reservation is left stranded. It does not appear in your usage history, because no model was called.

Every /v1 response from a rate-limited key — success or 429 — also carries:

HeaderMeaning
X-RateLimit-LimitRequests per minute allowed for this key.
X-RateLimit-RemainingRequests left in the current window. 0 on a 429.
X-RateLimit-ResetSeconds until the window resets — not an epoch timestamp. On a 429 this always matches Retry-After.

Use X-RateLimit-Remaining to pace yourself before you get refused, not just after.

Beta default: no per-key limit

Keys issued to date carry no rate limit, so in practice you will not see a 429 today. That is a beta convenience, not a guarantee. Handle 429 and Retry-After now rather than discovering the limit the day it is switched on.

Rate limits are separate from your prepaid balance. A 429 says "too fast"; a 402 says "not enough balance". See Costs.

Other current constraints

  • response_format / structured outputs is not supported. Both json_schema and json_object are silently dropped today — the request succeeds, but you get free-text markdown back, not the schema-conformant output you asked for, and nothing in the response signals that it happened. Do not build on this until it is confirmed fixed.
  • One key class. Only mb_live_ keys exist. There is no separate admin or management key, and no public management API. Key, balance, and usage administration happens in the console.
  • usage may be null. When the serving provider reports no usage, the field is present and null rather than omitted. Guard it.
  • Undocumented response members are not contract. Responses carry model (the serving tier) and a ModelBeat extra_fields object, of which only the members the reference documents are contract. Ignore the rest: they are unstable, may be redacted or removed without notice, and carry no correctness guarantee.
  • No published SLA. Private beta. See Support.

On this page