Skip to content

Mango Inference Docs

Documentation for Mango Inference

MangoBoost's platform for serving AI/ML workloads. Install it, deploy your first model, and run it reliably in production. Start with the quickstart.

Mango Inference is MangoBoost's inference cloud: open large language models, served behind an OpenAI-compatible HTTP API, billed per token. It is for developers and teams who want to call a model rather than operate one. Point an existing OpenAI or Anthropic client at a new base URL, name a model, and you are running. There is no cluster to size, no GPU to reserve, and nothing to pay for between requests.

Explore the docs

  • Getting Started


    Make your first API call in a few minutes.

    Quickstart

  • Core Concepts


    Understand how Mango Inference works, and its OpenAI-compatible API.

    Platform concepts

  • Capabilities


    Use tool calling, structured output, and reasoning.

    Tool calling

  • Reference


    Look up exact endpoints, parameters, and pricing.

    API reference

  • Platform


    Manage API keys, billing, credits, and usage.

    API keys

  • Resources


    Release notes and answers to common questions.

    FAQ