Skip to content

Model APIs

Mango Inference Model APIs are serverless endpoints for running LLMs. You send a request naming a model; Mango Inference runs it and returns the result. There are no servers, replicas, or scaling to manage. You pay per use.

Serverless model serving

  • No infrastructure. Call a model by ID; Mango Inference handles capacity.
  • OpenAI-compatible. Use the same request shape as the OpenAI API. See OpenAI compatibility.
  • Anthropic-compatible. Also send requests in the Anthropic Messages API shape, authenticated with the x-api-key header. See Messages reference.
  • Pay per use. Usage draws down your credits/balance at the rates in Models and pricing.

The models

Mango Inference serves open-weight text generation models, addressed by their upstream repository ID (vendor prefix included, case-sensitive). For example, zai-org/GLM-5.3. OpenAI and Anthropic model names (gpt-4o, claude-sonnet-4) do not resolve here, even on the endpoint that speaks their request shape; every request must name a Mango Inference model ID.

The catalogue is a live list, not a table in these docs

Which models are available changes as models are added and retired, so the authority is the platform rather than this page:

curl https://api.mangoboost.io/v1/models \
  -H "Authorization: Bearer $MANGOINFERENCE_API_KEY"

Returns the OpenAI-compatible model list: one entry per model, keyed by the id you pass as model.

Two things about that list are worth internalising, because they explain most "why can't I call this model?" surprises:

  • The list is scoped to your key. GET /v1/models returns the models that key may call, not a global catalogue. A key with no access to a model cannot call it, and will not see it.
  • A model missing from the list is a request, not a bug. Ask in Discord for a model you want deployed. See Platform support.

What a model entry gives you

The model list is deliberately thin. It is the OpenAI /v1/models schema, so an entry gives you the id to pass and little more. It is not a model card: it carries no context length, no capability flags, no price, and no quality metrics. That means the model entry answers "can I call this?", and the model's own upstream model card answers "should I, and how?". Use both:

Question Where the answer is
Which models can this key call? GET /v1/models, or Models in the console.
What exactly do I pass as model? The id field, copied verbatim.
Context length, training mix, intended use, benchmark scores The model's upstream model card.
Does it support tool calling / structured output / reasoning? The model card, then a real test call. See below.
What does it cost? Not published per model yet. See Models and pricing.
Which weights served my request? metadata.weight_version on the response.

Capabilities are properties of the model

Mango Inference passes capability parameters through to the model server rather than validating them, so a capability works if the model implements it. The platform side is uniform: every model reachable through these endpoints accepts the same request shape. The variation is entirely in the model:

Capability How you invoke it Where it varies
Streaming stream: true Platform-level; works on both endpoint families.
Tool calling tools (+ tool_choice) Per model. An unsupporting model typically ignores tools and answers in prose rather than erroring.
Structured output response_format Per model and its serving stack. json_object and json_schema are the modes these docs cover.
Reasoning Nothing (on by default where supported) Per model. Look for reasoning_content and usage.reasoning_tokens in the response.

Which specific models support what is not published

These docs deliberately do not carry a model-by-model capability matrix. The catalogue moves, and a stale "yes" in a table here is worse than no table: the failure mode is silent. A model that does not honour tools answers in prose, and a response_format that is not enforced returns plausible free-form text. The first thing that notices is json.loads in production.

So: check the model's own model card, then verify with one real call before you depend on a capability. For reasoning, the definitive test is whether reasoning_content comes back.

Serving path

There is one serving path: shared serverless endpoints. Every account calls the same on-demand capacity, and requests are priced per token with nothing reserved between them. There are no dedicated deployments, reserved capacity, or per-account replicas to choose from, so there is no sizing decision to make before your first call.

If your workload needs isolated or reserved capacity (a latency floor, a committed throughput, specific hardware), that is a conversation rather than a setting: raise it through Platform support.

Limits

Requests are subject to rate limits.

The surface is narrow by design: chat generation and the model list. There is no embeddings, batch, files, images, or audio endpoint. Those paths return 404. See OpenAI compatibility for the full boundary.