OpenAI compatibility¶
Mango Inference implements an OpenAI-compatible API. If you already use the OpenAI API or SDKs, you can point them at Mango Inference by changing two things: the base URL and the API key.
What this means¶
- Existing OpenAI client code works with minimal changes.
- Requests and responses follow the OpenAI schema.
- Tooling built on the OpenAI API (SDKs, agents, frameworks) can target Mango Inference. See OpenAI SDK.
Point a client at Mango Inference¶
from openai import OpenAI
client = OpenAI(
base_url="https://api.mangoboost.io/v1", # Mango Inference base URL
api_key="<your-api-key>",
)
Supported endpoints¶
Compatibility is deliberately narrow: Mango Inference serves chat, and the model list that goes with it. It is not a drop-in for the whole OpenAI platform.
| OpenAI endpoint | Supported | Notes |
|---|---|---|
POST /v1/chat/completions |
Yes | The primary surface. See reference. |
GET /v1/models |
Yes | Lists the models your key can call. |
POST /v1/completions (legacy) |
No | Use /v1/chat/completions. |
POST /v1/embeddings |
No | |
POST /v1/responses |
No | Use /v1/chat/completions. |
POST /v1/moderations |
No | |
POST /v1/images/*, /v1/audio/* |
No | Text generation only. |
/v1/files, /v1/batches |
No | No batch or file surface. |
Unsupported paths return HTTP 404, not a structured API error, so an SDK
method that targets one will usually surface as a "not found" rather than a
clear "unsupported" message. That is the fastest way to tell the difference
between a feature Mango Inference lacks and a request you got wrong.
Alongside these, Mango Inference serves POST /v1/messages and
POST /v1/messages/count_tokens in the Anthropic Messages shape. They are
not part of the OpenAI surface. They exist so Anthropic-ecosystem tools work
too. See Messages.
Known differences¶
Where Mango Inference differs from OpenAI, it adds rather than changes:
- Extra response fields. Responses may carry
reasoning_content,matched_stop, andmetadata.weight_version, none of which OpenAI returns. SDKs ignore unknown fields, so this is safe for normal client code. Strict schema validation on your side will still need to allow them. See Chat Completions. usage.reasoning_tokens. Reasoning models report reasoning tokens separately insidecompletion_tokens. See Reasoning.- Model IDs are not OpenAI model names. They are upstream repository IDs
such as
zai-org/GLM-5.3. Every request must name one;gpt-4oand friends do not resolve. See Models and pricing. - Sampling defaults are the model's, not OpenAI's. If your code relies on
OpenAI's default
temperature, set it explicitly instead. - No organization, project, or admin APIs. Account management lives in the console, not the API. See Operations.