Skip to content

OpenAI compatibility

Mango Inference implements an OpenAI-compatible API. If you already use the OpenAI API or SDKs, you can point them at Mango Inference by changing two things: the base URL and the API key.

What this means

  • Existing OpenAI client code works with minimal changes.
  • Requests and responses follow the OpenAI schema.
  • Tooling built on the OpenAI API (SDKs, agents, frameworks) can target Mango Inference. See OpenAI SDK.

Point a client at Mango Inference

from openai import OpenAI

client = OpenAI(
    base_url="https://api.mangoboost.io/v1",  # Mango Inference base URL
    api_key="<your-api-key>",
)

Supported endpoints

Compatibility is deliberately narrow: Mango Inference serves chat, and the model list that goes with it. It is not a drop-in for the whole OpenAI platform.

OpenAI endpoint Supported Notes
POST /v1/chat/completions Yes The primary surface. See reference.
GET /v1/models Yes Lists the models your key can call.
POST /v1/completions (legacy) No Use /v1/chat/completions.
POST /v1/embeddings No
POST /v1/responses No Use /v1/chat/completions.
POST /v1/moderations No
POST /v1/images/*, /v1/audio/* No Text generation only.
/v1/files, /v1/batches No No batch or file surface.

Unsupported paths return HTTP 404, not a structured API error, so an SDK method that targets one will usually surface as a "not found" rather than a clear "unsupported" message. That is the fastest way to tell the difference between a feature Mango Inference lacks and a request you got wrong.

Alongside these, Mango Inference serves POST /v1/messages and POST /v1/messages/count_tokens in the Anthropic Messages shape. They are not part of the OpenAI surface. They exist so Anthropic-ecosystem tools work too. See Messages.

Known differences

Where Mango Inference differs from OpenAI, it adds rather than changes:

  • Extra response fields. Responses may carry reasoning_content, matched_stop, and metadata.weight_version, none of which OpenAI returns. SDKs ignore unknown fields, so this is safe for normal client code. Strict schema validation on your side will still need to allow them. See Chat Completions.
  • usage.reasoning_tokens. Reasoning models report reasoning tokens separately inside completion_tokens. See Reasoning.
  • Model IDs are not OpenAI model names. They are upstream repository IDs such as zai-org/GLM-5.3. Every request must name one; gpt-4o and friends do not resolve. See Models and pricing.
  • Sampling defaults are the model's, not OpenAI's. If your code relies on OpenAI's default temperature, set it explicitly instead.
  • No organization, project, or admin APIs. Account management lives in the console, not the API. See Operations.