Model APIs¶
Mango Inference Model APIs are serverless endpoints for running LLMs. You send a request naming a model; Mango Inference runs it and returns the result. There are no servers, replicas, or scaling to manage. You pay per use.
Serverless model serving¶
- No infrastructure. Call a model by ID; Mango Inference handles capacity.
- OpenAI-compatible. Use the same request shape as the OpenAI API. See OpenAI compatibility.
- Anthropic-compatible. Also send requests in the Anthropic Messages API
shape, authenticated with the
x-api-keyheader. See Messages reference. - Pay per use. Usage draws down your credits/balance at the rates in Models and pricing.
The models¶
Mango Inference serves open-weight text generation models, addressed by their
upstream repository ID (vendor prefix included, case-sensitive). For example,
zai-org/GLM-5.3. OpenAI and Anthropic model names (gpt-4o,
claude-sonnet-4) do not resolve here, even on the endpoint that speaks their
request shape; every request must name a Mango Inference model ID.
The catalogue is a live list, not a table in these docs¶
Which models are available changes as models are added and retired, so the authority is the platform rather than this page:
curl https://api.mangoboost.io/v1/models \
-H "Authorization: Bearer $MANGOINFERENCE_API_KEY"
Returns the OpenAI-compatible model list: one entry per model, keyed by the
id you pass as model.
Sign in at https://inference.mangoboost.io and open Models (https://inference.mangoboost.io/models).
Two things about that list are worth internalising, because they explain most "why can't I call this model?" surprises:
- The list is scoped to your key.
GET /v1/modelsreturns the models that key may call, not a global catalogue. A key with no access to a model cannot call it, and will not see it. - A model missing from the list is a request, not a bug. Ask in Discord for a model you want deployed. See Platform support.
What a model entry gives you¶
The model list is deliberately thin. It is the OpenAI /v1/models schema, so
an entry gives you the id to pass and little more. It is not a model card: it
carries no context length, no capability flags, no price, and no quality
metrics. That means the model entry answers "can I call this?", and the
model's own upstream model card answers "should I, and how?". Use both:
| Question | Where the answer is |
|---|---|
| Which models can this key call? | GET /v1/models, or Models in the console. |
What exactly do I pass as model? |
The id field, copied verbatim. |
| Context length, training mix, intended use, benchmark scores | The model's upstream model card. |
| Does it support tool calling / structured output / reasoning? | The model card, then a real test call. See below. |
| What does it cost? | Not published per model yet. See Models and pricing. |
| Which weights served my request? | metadata.weight_version on the response. |
Capabilities are properties of the model¶
Mango Inference passes capability parameters through to the model server rather than validating them, so a capability works if the model implements it. The platform side is uniform: every model reachable through these endpoints accepts the same request shape. The variation is entirely in the model:
| Capability | How you invoke it | Where it varies |
|---|---|---|
| Streaming | stream: true |
Platform-level; works on both endpoint families. |
| Tool calling | tools (+ tool_choice) |
Per model. An unsupporting model typically ignores tools and answers in prose rather than erroring. |
| Structured output | response_format |
Per model and its serving stack. json_object and json_schema are the modes these docs cover. |
| Reasoning | Nothing (on by default where supported) | Per model. Look for reasoning_content and usage.reasoning_tokens in the response. |
Which specific models support what is not published
These docs deliberately do not carry a model-by-model capability matrix.
The catalogue moves, and a stale "yes" in a table here is worse than no
table: the failure mode is silent. A model that does not honour tools
answers in prose, and a response_format that is not enforced returns
plausible free-form text. The first thing that notices is json.loads in
production.
So: check the model's own model card, then verify with one real call before
you depend on a capability. For reasoning, the definitive test is whether
reasoning_content comes back.
Serving path¶
There is one serving path: shared serverless endpoints. Every account calls the same on-demand capacity, and requests are priced per token with nothing reserved between them. There are no dedicated deployments, reserved capacity, or per-account replicas to choose from, so there is no sizing decision to make before your first call.
If your workload needs isolated or reserved capacity (a latency floor, a committed throughput, specific hardware), that is a conversation rather than a setting: raise it through Platform support.
Limits¶
Requests are subject to rate limits.
The surface is narrow by design: chat generation and the model list. There is no
embeddings, batch, files, images, or audio endpoint. Those paths return 404.
See OpenAI compatibility for the
full boundary.