baseURL: https://platform.belvedir.ai/api/v1/route with your bv_live_ key and chat.completions.create lands here. Belvedir picks the model, runs the call, and returns the provider’s response unchanged. The two-line client change is on Route your calls; how routers, tiers, Smart routing and Automatic updates decide is on Model Routing.
Request
- Authorization: Bearer token with your API key, or use the
x-api-keyheader. - Body: a standard OpenAI chat-completions body:
messages,tools,tool_choice,response_format, images in content parts,stream, sampling parameters. Bodies cap at 1 MB (413beyond it). - Session: pass your session id in the
x-session-idheader or asession_idbody field (stripped before forwarding) to pin the conversation to one model for up to 24 hours. Pinning is described under session pinning.
Model ids
model works two ways:
- A real model id (
x-ai/grok-4.6,anthropic/claude-fable-5, a baregpt-5.2; the catalog with identifiers is on Available Models): the call goes through a router anchored on that model, created the first time you name it and listed on the Routers page as “From your code”. The named model is the ceiling and answers every conversational call; cheaper tiers take only the machine-shaped tasks underneath it. Code that names Grok in one place and Claude in another keeps working and ends up with one router per model. Deleting one means the next call naming that model recreates it. A project that deleted all its routers forwards an explicitmodelunchanged (x-belvedir-tier: client). "auto"(or nomodel): the call goes to the project’s auto router. On a project with no router it returns a clear400.
claude-sonnet-5, grok-4.6), so a call site migrated from another SDK keeps the id it already names. OpenAI ids that only run on the Responses API (the gpt-*-pro and codex families) are refused with a 400 naming the base chat model (or a chat-latest variant) to use instead.
Where the call runs
Hosted models run on Belvedir’s provider accounts. Hosted open models run on the fastest available provider that meets Belvedir’s data-retention requirements; that can be a pricier host than the cheapest one, andusage.cost reflects what actually served. To choose differently, pass a provider preference object on the body (for example {"sort": "price"}); it replaces the default and is forwarded as is. Client-side routing overrides (models, route, plugins, transforms, web_search_options) are ignored: the routed model is the only routing input.
A route to a local model (Ollama-style name:tag ids) or to one of your fine-tuned models returns a clear 400 here, unless you registered your own deployment for that model id under Cloud Inference: then the call runs on your endpoint with your provider key, metered but never billed. Without a registered endpoint, use the decision endpoint and run the call yourself. Endpoints you register must be https, publicly resolvable, and not Belvedir’s own hosts; a 3xx from them is not followed and returns 502 with code: "gateway_redirect", so point the base URL at the final host. How to register one is on Cloud Inference.
Streaming, tools, structured output
- Streaming:
stream: truereturns the provider’s SSE byte for byte, with the usage block on the final chunk before[DONE](Belvedir turns onstream_options.include_usageupstream, so you don’t have to). A streamed response the client aborts is billed on an estimate of what was streamed. - Tools:
toolsandtool_choicepass through unchanged. Claude Fable 5.1 and Claude Opus 5.5 reject forced tool choice; see Capabilities and Limits for how Belvedir handles that. - Structured output:
response_format(json_object,json_schema) passes through. With Automatic updates on, a non-streaming machine-shaped request with a strict output contract can be served by the Small tier first and escalated when the answer fails validation; the response says which happened inx-belvedir-cascade: served-smallorescalated, and both attempts meter. The toggle is described on Model Routing.
Response headers
Which model served the call comes back inmodel on the response body and in these headers:
Token counts and cost
Every response carries the standard OpenAIusage block: prompt_tokens, completion_tokens, total_tokens, plus whatever detail the provider reports (for example prompt_tokens_details.cached_tokens). Non-streamed calls carry it on the body; streamed calls get it on the final chunk. When Token compression ran, the counts are what the model actually received.
What the call cost comes back as usage.cost (USD, a number) on the body and in the x-belvedir-cost header; streamed calls carry usage.cost on the final usage chunk. It is the amount billed to your organization for that call: the model’s per-token rate from the Pricing page (Inference → Pricing), or the provider’s reported cost for models billed that way, so summing usage.cost across calls reproduces your usage page. Cached input is priced at the cache-read rate. Calls served from your own registered endpoint report 0. A model with no known price omits the field, and the usage page shows it unpriced. When a server-side retry fired, the metered first attempt also bills, and usage.cost covers only the answer you received.
Calls are attributed to the calling API key and draw on your organization’s balance. An organization with no balance and no card gets 402 before any model is called; balance, Auto reload and pay as you go are set under Organization Settings → Billing.
Token compression
Under Token compression (a per-project toggle under Project Permissions, on by default), long prompts are shortened before the call; very short prompts are sent as is. The routing decision always reads your original text, and the model’s response is never compressed. Tokens saved come back inx-belvedir-tokens-saved and count toward the savings shown on your Home and Billing pages. If compression is ever unavailable, your original prompt is sent unchanged; a routed call never fails because of it. Turn it off when every word in your prompts carries weight: lab protocols, legal or medical text, dense technical specs.
Reasoning controls
When a call lands on a reasoning model and itsmax_tokens (or max_completion_tokens) is under 2,048, or it routed to the Small tier, Belvedir turns thinking off for that call so the budget goes to the answer instead of an unfinished reasoning trace. Set reasoning, reasoning_effort, or chat_template_kwargs yourself to override. Your setting is honored in translated form: each provider family accepts a different reasoning vocabulary, so Belvedir maps your value to what the serving model accepts, and an off setting (none, minimal, enable_thinking: false, thinking: {"type": "disabled"}) becomes that model’s working off-switch. Graded effort levels are honored on OpenAI reasoning models and on open models that support them; Claude models on this endpoint honor only off, Claude Fable always thinks, and Grok’s reasoning is fixed per model id, so reasoning controls are dropped for it.
Server-side retries
Belvedir retries a doomed call rather than bill you for an empty answer. When a retry fired, the response says so inx-belvedir-retry:
thinking-disabled: a reasoning model spent the whole completion budget on its thinking trace and returned an empty answer. Belvedir retries once, same model, thinking off. This fires even when your own reasoning setting caused it.content-filter-fallback: a model’s safety filter returned an empty answer (finish_reason: "content_filter") for a prompt the router chose it for. Belvedir retries on your router’s other tiers, cheapest suitable first. Only when the router picked the model: if your code named it, you get that model’s verdict unretried.tool-choice-auto: the serving model rejects forced tool choice, so the call was sent withtool_choice: "auto"; see Capabilities and Limits.
Errors
Upstream provider errors are returned with the provider’s status and body. The full limits table is on Capabilities and Limits.