POST /api/v1/messages serves the Messages API, untranslated: point the official Anthropic SDK at Belvedir and requests pass through to Anthropic byte-faithfully.
authToken option (or the ANTHROPIC_AUTH_TOKEN env var) works as the key carrier too; Belvedir reads the key from either Authorization: Bearer or x-api-key.
What passes through intact
- Prompt caching:
cache_controlbreakpoints (per-block and top-level) reach Anthropic unmodified, and cache reads bill at the cache-read rate. Verify fidelity yourself by readingusage.cache_read_input_tokensoff responses, exactly as you would against Anthropic directly. Cache writes bill at Anthropic’s published write premiums (1.25× for 5-minute, 2× for 1-hour). - Thinking: adaptive thinking,
displaysettings, thinking blocks in responses and streams; Belvedir never injects or strips a thinking config on this surface. anthropic-betaheaders: forwarded verbatim, exceptoauth-*betas (a client-auth artifact), which are dropped. Fast mode, compaction, context management and future betas work without waiting on Belvedir; fast-mode calls bill at Anthropic’s published fast rate.- Streaming: Anthropic’s native event shape (
message_start,content_block_delta,message_delta), byte for byte. A frontend that parses Anthropic SSE keeps working unchanged. - The Claude Agent SDK: it honors the same env vars for all its traffic, subagents included. Set
ANTHROPIC_BASE_URL=https://platform.belvedir.ai/apiandANTHROPIC_API_KEY=<bv_live_ key>(orANTHROPIC_AUTH_TOKEN) and its inference goes through Belvedir.
Models
This surface is model-pinned by design: pass a real model id."auto" and routing live on the OpenAI-compatible endpoint. Accepted ids:
- Claude models, bare (
claude-sonnet-5) or prefixed (anthropic/claude-sonnet-5). - These hosted open models, which Belvedir serves behind an Anthropic-compatible API:
zai-org/GLM-5.2-FP8,moonshotai/Kimi-K2.6,deepseek/deepseek-v4-flash-0731,openai/gpt-oss-120b,google/gemma-4-31B-it,nvidia/Gemma-4-31B-IT-NVFP4.
400 pointing at the OpenAI-compatible endpoint. A Chinese-lab model on a project with Chinese models off returns 403 (see Data handling).
Server tools
Web search and web fetch work here; other Anthropic server tools do not.web_search_<date> and web_fetch_<date> tools pass through and are metered from the response’s usage.server_tool_use: each web search bills a flat 1 cent on top of tokens (Anthropic’s list price of $10 per 1,000 searches); web fetch has no per-call fee. Code execution, computer use, the text editor / bash / memory / tool-search tools, the MCP connector (mcp_servers) and containers bill by container time or outside the usage object, so a request carrying one returns 400 before anything runs. Custom tools (a name plus input_schema, with no type or type: "custom") work everywhere.
Forced tool choice is not translated on this surface: Claude Fable 5.1 and Claude Opus 5.5 reject it, so send {"type": "auto"} yourself (see Capabilities and Limits).
Cost and metering
Every token is metered from the native usage object, cache writes and fast mode included, at the rates on the Pricing page (Inference → Pricing), attributed to the calling API key. On non-streamed calls, what the call cost rides thex-belvedir-cost response header; the body is never rewritten. Streamed calls carry no cost header: read spend off the usage page. A streamed response the client aborts is billed on an estimate of what was streamed.
Errors and limits
Errors come back in Anthropic’s error envelope, so the SDK’s typed errors work; upstream errors pass through verbatim. Request bodies cap at 20 MB (413 beyond it). The key auth, the 402 spend gate and the rate limit (25 requests per second per API key, burst 300, adjustable per organization) are the same as on chat completions.
Message Batches
With the Anthropic client pointed athttps://platform.belvedir.ai/api, client.messages.batches.create / retrieve / results / cancel / list / delete run against /api/v1/messages/batches, which speaks Anthropic’s Message Batches wire shapes over Belvedir’s batch tier. It is documented with the other batch endpoints on Batches.