The surface is narrow on purpose: two metered endpoints and model: "auto". Model identity is not exposed.
ModelBeatbeta

Routing

What auto does, the intelligence tiers, the gates a request has to clear, what happens on a fallback, and why model identity is not on the response.

model takes auto or one of three intelligence tiers. You do not name a model, and the response does not name one back. ModelBeat picks the cheapest model that can actually serve the request, executes it, and charges you what it cost.

r = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "summarise this changelog"}],
    max_tokens=100,
)

Intelligence tiers

auto is the default and the one most callers want. When you need a floor on capability rather than on price, name a tier instead:

modelWhat it means
autoCheapest candidate that clears the request's gates.
modelbeat-fastFavours latency and price. The widest pool.
modelbeat-standardA middle capability floor.
modelbeat-advancedThe highest floor — only the strongest candidates qualify.

A tier raises the quality floor a candidate must clear before cost decides among what survives; it does not pin anything. Everything below still applies — the gates, the fallback behaviour, and the fact that nothing names a model.

Tiers name a capability band, not a model. That is deliberate and it is why they exist alongside the position in the rest of this page: the pool behind a tier changes without notice, and nothing you build should assume otherwise.

The gates

Selection is not "cheapest thing in the catalogue". A candidate has to clear four hard gates before cost is allowed to decide:

  • Output cap. A model whose maximum output length is below your max_tokens is excluded.
  • Context window. A model whose input limit is below your prompt is excluded.
  • Capabilities. An image in the request excludes text-only candidates. The same applies to JSON mode and function calling.
  • Quality floor. A code or reasoning task hint requires a candidate above a minimum quality tier.

Among what survives, cost decides. That is why a small request lands on the cheapest rung, and a 32,000-token output request does not: the cheap rungs cap output well below the top of the range, and a large max_tokens gates them out. Raising max_tokens "just in case" is the single most common reason a request costs more than you expected — and it also inflates the balance hold. See Costs.

If nothing survives the gates you get a 400 with a reason such as no_eligible_candidate. That is a statement about your request, so retrying it unchanged will fail again. See Errors.

Fallbacks

If the chosen model fails, ModelBeat retries down a short fallback chain rather than failing your request. This is invisible to you by design: you get a completion, and the hold placed against your balance covers the chosen candidate and its fallbacks, so a fallback never costs more than was reserved.

A fallback is not an error. The request succeeded.

Why the served model is not on the response

Not exposed, and not selectable

During private beta, the identity of the model that served a request is not part of the published API. There is no field to read it from, no way to pin, prefer, or exclude a specific model, and the only catalogue GET /v1/models lists is the three tiers above. What a response does tell you is the tier that served it — never which model.

The reason is that the routing pool is not a stable contract yet. Models enter and leave it, prices move, and the gates are still being tuned. Anything built against a specific model id today would break, quietly, on a catalogue change you never see. Holding the surface at auto keeps the one thing that is stable — an OpenAI-shaped completion with an honest price attached — from being coupled to a moving part.

What you do get back is the tier. The response's model carries the serving tier rather than a model id, and extra_fields.routing_info carries level (the same tier, short form) and is_fallback. Nothing below the tier — provider, model, key — is returned.

Everything else under extra_fields is unstable, may be redacted or removed without notice, and carries no correctness guarantee. Read the reference; ignore the rest. See Limits.

If you need a per-request record

ModelBeat writes an immutable decision record for every request — the routing reason, the routing level, whether a fallback occurred, how many attempts it took, tokens, cost, latency, and timestamp. That record is retained regardless of what the API response carries, and it is what an EU AI Act or internal-audit enquiry is answered from.

It is reachable, but not from this API:

  • Support. Quote the X-Request-Id from the response and we can pull the exact call, served model included. This is the route to use during private beta. See Errors.
  • The console. Per-request history for your tenant, filterable, at app.modelbeat.ai. During beta it shows the tier that served each request rather than the model, so for the model itself, ask support.

If your obligations require the served model to be readable programmatically, on the response, tell us. It is a deliberate beta constraint rather than a permanent one, and it is prioritised by demand.

On this page