Skip to main content
These are the models Belvedir’s inference endpoints can execute, with the identifier you pass as model (or pick in a router on the Routers page). This is the current catalog; the Pricing page (Inference → Pricing) is authoritative and carries the per-token rates, so prices are not repeated here. Ids not on it bill at the provider’s reported cost for that call.

Identifier spelling

Belvedir uses provider-prefixed identifiers for every provider except OpenAI, whose bare gpt-* ids are the convention. Bare ids from provider-native SDKs are normalized server-side (claude-sonnet-5 → anthropic/claude-sonnet-5, grok-4.6 → x-ai/grok-4.6, gemini-* and gemma-* → google/*, mistral-*, mixtral-*, and codestral-* → mistralai/*, deepseek-* → deepseek/*), so a call site migrated from another SDK keeps working with the id it already names.

Anthropic

Served on Anthropic’s API. Claude models are also servable through the Anthropic-native Messages passthrough, which keeps prompt caching, thinking, beta headers, and native streaming intact. Dated snapshot ids (claude-opus-4-5-20251101, …) resolve to the same rates. Claude Fable 5.1 and Claude Opus 5.5 do not accept forced tool choice and never run with thinking off; how Belvedir handles both is on Capabilities and Limits.

OpenAI

Served on OpenAI’s API. Bare ids are the norm; the openai/ prefix is also accepted. Any current OpenAI chat id (gpt-*, chatgpt-*, o-series) routes here; the priced line-up, newest first: Dated snapshot ids and the -chat-latest variants resolve to the same rates. GPT-6 Astra accepts function tools only on OpenAI’s Responses API, so Belvedir sends Astra tool calls there and returns the usual chat-completions shape (JSON or streamed); your OpenAI client code does not change. OpenAI’s pro and codex models are served only on their Responses API, which the router doesn’t speak; they return a clear error naming the base chat model (or a chat-latest variant) to use instead.

xAI

Served on xAI’s API and billed at the provider’s reported cost per call. Any model xAI serves works; the current line-up:

Open models

Hosted open models run on the fastest available provider that meets Belvedir’s data-retention requirements. The Batch tier column says whether the id is accepted by batches. Beyond this list, any open-model id in the org/model spelling works as a pass-through on chat completions and bills at the provider’s reported cost per call. Pass-through ids have no batch tier. With Chinese models off under Project Permissions, Chinese-lab identifiers are unavailable for that project across the whole catalog, pass-through ids included; see Data handling.

Your own models

  • Local models (Ollama-style name:tag ids such as qwen3:4b, llama3.1:8b, gpt-oss:20b): the decision endpoint can route to them, but your app executes the call; the inference endpoints return a clear 400 for them.
  • Fine-tuned models (your project’s trained models, by name): same rule; the decision endpoint routes to them and your app serves them.
  • Bring your own provider: register a model id with your own base URL and key under Cloud Inference in the dashboard, and the inference endpoint executes that id on your endpoint instead, metered but never billed. This lifts the 400 for local and fine-tuned ids too. See Cloud Inference.