bv_live_ API key. This page is the summary of what each one can and cannot do, and the one place every hard limit is listed. The request and response contracts live in the API reference, linked per surface below.
The four surfaces
A fifth endpoint,
POST /api/v1/route, makes the routing decision only and executes nothing: use it when your own code runs the model (local models, your own provider accounts).
What each surface supports
What Smart routing and Automatic updates change is on Model Routing.
Models each surface executes
The full catalog with identifiers is on Available Models. In summary:- Chat completions: every Claude model, every callable OpenAI chat model, Grok, and the hosted open models. Model ids you registered under Cloud Inference execute on your own endpoint (metered, never billed).
- Messages passthrough: Claude models, plus the six hosted open models listed on Messages. Nothing else; other ids get a
400pointing at the OpenAI-compatible endpoint. - Embeddings: any OpenAI
text-embedding-*id is accepted, but an unpriced one returns400; the priced ones aretext-embedding-3-small,text-embedding-3-large, andtext-embedding-ada-002. - Batches: Anthropic models (up to 100,000 requests per batch), OpenAI models (up to 50,000), and the three hosted open models with a batch tier: GLM 5.2, Kimi K2.6, and Gemma 4 31B (up to 1,000). One batch can mix providers.
What Belvedir cannot do
Each of these fails fast with a clear error rather than an opaque upstream failure:- Local models (Ollama-style
name:tagids) and your project’s fine-tuned models are never executed by the router:400. Use the decision endpoint and run the call yourself, or register a deployment under Cloud Inference, which lifts the restriction. - OpenAI’s Responses API: the router speaks chat completions only, so OpenAI’s pro and codex families return
400naming the base chat model to use instead. - Anthropic Managed Agents / Agent API sessions run their inference inside Anthropic’s own orchestration; there is no base URL to override, for Belvedir or anyone else. Instrument those workloads with the Belvedir SDK for observability instead.
- No routing on the Messages passthrough or in batches: both surfaces are model-pinned;
"auto"returns400. Routing lives on the chat-completions endpoint. - Forced tool choice on Claude Fable 5.1 and Claude Opus 5.5: both reject
tool_choice: "required"and named-function choices (Anthropic returns400). On chat completions and in batches Belvedir sends those calls withtool_choice: "auto"instead; only chat completions says so on the wire (x-belvedir-retry: tool-choice-auto), while a batch result carries no marker, so check batch results for a text answer where you expected a tool call. Name the tool in your prompt and do not assume atool_callsfinish. The Messages passthrough is untranslated: send{"type": "auto"}yourself there. Neither model runs with thinking off, so an off setting has no effect on them. - Anthropic server tools: web search and web fetch only, on the Messages passthrough only. Other server tools return
400; native batchparamsrefuse every server tool. Details on Messages. - No redirects from your endpoint: a registered endpoint that answers
3xxreturns502(gateway_redirect) instead of being followed; register the final host. Base URLs must behttps, public, and not Belvedir’s own hosts. - Unpriced Anthropic, OpenAI and hosted open model ids return
400. Grok and open-model pass-through ids bill at the provider’s reported cost per call, so they need no rate entry. - No batch tier for open-model pass-through ids or for DeepSeek V4 Flash, gpt-oss 120B and the NVFP4 Gemma variant.
- No streaming inside batches, and Anthropic-native
paramsbodies are not accepted for OpenAI models. - No token-counting, files, image generation, audio, or realtime APIs: the surfaces above are text-in, text-out chat, embeddings, and batches.
- Chinese-lab models when the project turned them off (Project Permissions):
403on every surface, never a silent substitution. See Data handling.
Reliability backstops
Belvedir retries doomed calls (an empty answer after a reasoning trace ate the budget, a spurious content-filter refusal on a model the router chose, a rejected thinking-off switch) rather than bill you for nothing, and says so in thex-belvedir-retry header. What each retry does and what it meters is on chat completions.
Operational limits
Beyond the rate limits the endpoints return
429; retry after a short backoff. When batch capacity is momentarily full, submission returns a retryable 503; retry after the Retry-After header.