Capabilities
What actually works today beyond the base request shape — tool calling, vision, and the in-repo SDKs — each confirmed by live testing, not assumed.
Why this page exists
The base contract in the reference covers model, messages,
max_tokens, temperature, top_p, stop, and stream. It says nothing about the
request shapes below, so there was no way to tell whether they worked. Each one here
was confirmed against a real deployed edge, not inferred from reading code — see
Limits for what was tested and found broken instead.
Tool calling
Tool calling is a first-class, router-aware capability: the router only serves a tools-enabled request from a model that actually supports function calling, and this is covered by the router's own test suite. It is not a "probably passes through" extra.
r = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "What's the weather in Helena, Montana?"}],
tools=[{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}],
tool_choice="auto",
)tool_choice supports the standard values:
| Value | Behavior |
|---|---|
"auto" | Model decides whether to call a tool. |
"required" | Forces a call. |
"none" | Suppresses tool use — the model answers in text only. |
{"type": "function", "function": {"name": "..."}} | Forces a specific named function. |
tool_choice: none — know the history
Live testing on 2026-08-28 found tool_choice: "none" silently ignored on both local
and staging — the model called the tool anyway, with no signal in the response that
the constraint was dropped. That defect (and a related gap in named-function forcing
on some model families) was root-caused and fixed
(PR #1029). If you built a
workaround around "none" not working, you can remove it.
Vision
Image content parts work. Send an image as a data URI or URL in a content array, the
same shape as OpenAI's vision API:
r = client.chat.completions.create(
model="auto",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What color is this?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}},
],
}],
)If you test this with a tiny image, read this first
A 1×1-pixel test image produced a wrong-looking answer during verification — that turned out to be a degenerate-input artifact (the image was too small for the vision encoder to resolve), not a dropped feature. A 64×64 solid-color control image answered correctly. Test with a real image before concluding vision is broken.
Anthropic Messages and OpenAI Responses shapes
POST /v1/messages (Anthropic's wire format) and POST /v1/responses (OpenAI's
Responses wire format) are both supported, streaming and non-streaming, as of
2026-09-01.
Sending Anthropic's shape does not pin you to an Anthropic-served model: ModelBeat's own
catalog-driven routing still decides which model serves the request, the same as it does
for /v1/chat/completions.
import anthropic
client = anthropic.Anthropic(
api_key="mb_live_...",
base_url="https://api.modelbeat.ai/v1",
)
r = client.messages.create(
model="auto",
max_tokens=100,
messages=[{"role": "user", "content": "Write a haiku about prepaid inference."}],
)See Limits for the one billing caveat worth knowing before you rely on this: prompt-cache tokens are billed at the regular input rate, not Anthropic's cheaper cache-read rate.
SDKs
sdks/python and sdks/typescript are real, versioned clients — not internal tooling.
Both are thin wrappers around the official OpenAI SDK: you get typed routing_info
(tier + fallback flag; provider/model identity is None/undefined during the private
beta — see Limits) and typed cost
instead of untyped extra fields, plus a raw escape hatch to the underlying OpenAI
client for anything the wrapper doesn't cover yet.
pip install modelbeat # Python
npm install @modelbeat/sdk # TypeScriptBoth wrappers currently cover only chat.completions.create and completions.create.
For /v1/embeddings or /v1/feedback, use the raw client. Full usage, error handling,
and current limits are in each package's own README:
sdks/python/README.md,
sdks/typescript/README.md.
Not covered here
- Structured outputs (
response_format) — confirmed broken, not documented as working. See Limits. Idempotency-Key— supported: send the same key on a retry and the request is billed once. Reusing a key for a different body, or one whose reservation already settled, is refused with409. See the API reference.
Next
- Limits. What is supported, and what deliberately or currently is not.
- API reference. Full schemas, with a playground.