model (or pick in a router on the Routers page). This is the current catalog; the Pricing page (Inference → Pricing) is authoritative and carries the per-token rates, so prices are not repeated here. Ids not on it bill at the provider’s reported cost for that call.
Identifier spelling
Belvedir uses provider-prefixed identifiers for every provider except OpenAI, whose baregpt-* ids are the convention. Bare ids from provider-native SDKs are normalized server-side (claude-sonnet-5 → anthropic/claude-sonnet-5, grok-4.6 → x-ai/grok-4.6, gemini-* and gemma-* → google/*, mistral-*, mixtral-*, and codestral-* → mistralai/*, deepseek-* → deepseek/*), so a call site migrated from another SDK keeps working with the id it already names.
Anthropic
Served on Anthropic’s API. Claude models are also servable through the Anthropic-native Messages passthrough, which keeps prompt caching, thinking, beta headers, and native streaming intact.
Dated snapshot ids (
claude-opus-4-5-20251101, …) resolve to the same rates. Claude Fable 5.1 and Claude Opus 5.5 do not accept forced tool choice and never run with thinking off; how Belvedir handles both is on Capabilities and Limits.
OpenAI
Served on OpenAI’s API. Bare ids are the norm; theopenai/ prefix is also accepted. Any current OpenAI chat id (gpt-*, chatgpt-*, o-series) routes here; the priced line-up, newest first:
Dated snapshot ids and the
-chat-latest variants resolve to the same rates. GPT-6 Astra accepts function tools only on OpenAI’s Responses API, so Belvedir sends Astra tool calls there and returns the usual chat-completions shape (JSON or streamed); your OpenAI client code does not change. OpenAI’s pro and codex models are served only on their Responses API, which the router doesn’t speak; they return a clear error naming the base chat model (or a chat-latest variant) to use instead.
xAI
Served on xAI’s API and billed at the provider’s reported cost per call. Any model xAI serves works; the current line-up:Open models
Hosted open models run on the fastest available provider that meets Belvedir’s data-retention requirements. The Batch tier column says whether the id is accepted by batches.
Beyond this list, any open-model id in the
org/model spelling works as a pass-through on chat completions and bills at the provider’s reported cost per call. Pass-through ids have no batch tier.
With Chinese models off under Project Permissions, Chinese-lab identifiers are unavailable for that project across the whole catalog, pass-through ids included; see Data handling.
Your own models
- Local models (Ollama-style
name:tagids such asqwen3:4b,llama3.1:8b,gpt-oss:20b): the decision endpoint can route to them, but your app executes the call; the inference endpoints return a clear400for them. - Fine-tuned models (your project’s trained models, by name): same rule; the decision endpoint routes to them and your app serves them.
- Bring your own provider: register a model id with your own base URL and key under Cloud Inference in the dashboard, and the inference endpoint executes that id on your endpoint instead, metered but never billed. This lifts the
400for local and fine-tuned ids too. See Cloud Inference.