Skip to main content
Belvedir serves inference on four surfaces, all behind the same bv_live_ API key. This page is the summary of what each one can and cannot do, and the one place every hard limit is listed. The request and response contracts live in the API reference, linked per surface below.

The four surfaces

A fifth endpoint, POST /api/v1/route, makes the routing decision only and executes nothing: use it when your own code runs the model (local models, your own provider accounts).

What each surface supports

What Smart routing and Automatic updates change is on Model Routing.

Models each surface executes

The full catalog with identifiers is on Available Models. In summary:
  • Chat completions: every Claude model, every callable OpenAI chat model, Grok, and the hosted open models. Model ids you registered under Cloud Inference execute on your own endpoint (metered, never billed).
  • Messages passthrough: Claude models, plus the six hosted open models listed on Messages. Nothing else; other ids get a 400 pointing at the OpenAI-compatible endpoint.
  • Embeddings: any OpenAI text-embedding-* id is accepted, but an unpriced one returns 400; the priced ones are text-embedding-3-small, text-embedding-3-large, and text-embedding-ada-002.
  • Batches: Anthropic models (up to 100,000 requests per batch), OpenAI models (up to 50,000), and the three hosted open models with a batch tier: GLM 5.2, Kimi K2.6, and Gemma 4 31B (up to 1,000). One batch can mix providers.

What Belvedir cannot do

Each of these fails fast with a clear error rather than an opaque upstream failure:
  • Local models (Ollama-style name:tag ids) and your project’s fine-tuned models are never executed by the router: 400. Use the decision endpoint and run the call yourself, or register a deployment under Cloud Inference, which lifts the restriction.
  • OpenAI’s Responses API: the router speaks chat completions only, so OpenAI’s pro and codex families return 400 naming the base chat model to use instead.
  • Anthropic Managed Agents / Agent API sessions run their inference inside Anthropic’s own orchestration; there is no base URL to override, for Belvedir or anyone else. Instrument those workloads with the Belvedir SDK for observability instead.
  • No routing on the Messages passthrough or in batches: both surfaces are model-pinned; "auto" returns 400. Routing lives on the chat-completions endpoint.
  • Forced tool choice on Claude Fable 5.1 and Claude Opus 5.5: both reject tool_choice: "required" and named-function choices (Anthropic returns 400). On chat completions and in batches Belvedir sends those calls with tool_choice: "auto" instead; only chat completions says so on the wire (x-belvedir-retry: tool-choice-auto), while a batch result carries no marker, so check batch results for a text answer where you expected a tool call. Name the tool in your prompt and do not assume a tool_calls finish. The Messages passthrough is untranslated: send {"type": "auto"} yourself there. Neither model runs with thinking off, so an off setting has no effect on them.
  • Anthropic server tools: web search and web fetch only, on the Messages passthrough only. Other server tools return 400; native batch params refuse every server tool. Details on Messages.
  • No redirects from your endpoint: a registered endpoint that answers 3xx returns 502 (gateway_redirect) instead of being followed; register the final host. Base URLs must be https, public, and not Belvedir’s own hosts.
  • Unpriced Anthropic, OpenAI and hosted open model ids return 400. Grok and open-model pass-through ids bill at the provider’s reported cost per call, so they need no rate entry.
  • No batch tier for open-model pass-through ids or for DeepSeek V4 Flash, gpt-oss 120B and the NVFP4 Gemma variant.
  • No streaming inside batches, and Anthropic-native params bodies are not accepted for OpenAI models.
  • No token-counting, files, image generation, audio, or realtime APIs: the surfaces above are text-in, text-out chat, embeddings, and batches.
  • Chinese-lab models when the project turned them off (Project Permissions): 403 on every surface, never a silent substitution. See Data handling.

Reliability backstops

Belvedir retries doomed calls (an empty answer after a reasoning trace ate the budget, a spurious content-filter refusal on a model the router chose, a rejected thinking-off switch) rather than bill you for nothing, and says so in the x-belvedir-retry header. What each retry does and what it meters is on chat completions.

Operational limits

Beyond the rate limits the endpoints return 429; retry after a short backoff. When batch capacity is momentarily full, submission returns a retryable 503; retry after the Retry-After header.