embeddings.create against the router base URL runs OpenAI’s embedding models on Belvedir’s account with your bv_live_ key. The same client you use for chat completions works unchanged.
Request
- Authorization: Bearer token with your API key, or use the
x-api-keyheader. - Body: a standard OpenAI embeddings body (
model,input, and the optionaldimensionsandencoding_format). Bodies cap at 4 MB (413beyond it).
Models
Anytext-embedding-* id is accepted, with or without the openai/ prefix, but an unpriced one is refused with 400. The priced ids are text-embedding-3-small, text-embedding-3-large, and text-embedding-ada-002. Embeddings are model-pinned: there is nothing to route, and "auto" is not accepted.
Cost
Input tokens are metered like every routed call, at the rates on the Pricing page (Inference → Pricing), attributed to the calling API key. The cost of the call rides thex-belvedir-cost response header.
Limits
The402 spend gate and the rate limit (25 requests per second per API key, burst 300, adjustable per organization) are the same as on chat completions. The full table is on Capabilities and Limits.