Mango Inference Docs
Documentation for Mango Inference¶
MangoBoost's platform for serving AI/ML workloads. Install it, deploy your first model, and run it reliably in production. Start with the quickstart.
Mango Inference is MangoBoost's inference cloud: open large language models, served behind an OpenAI-compatible HTTP API, billed per token. It is for developers and teams who want to call a model rather than operate one. Point an existing OpenAI or Anthropic client at a new base URL, name a model, and you are running. There is no cluster to size, no GPU to reserve, and nothing to pay for between requests.
Explore the docs¶
-
Getting Started
Make your first API call in a few minutes.
-
Core Concepts
Understand how Mango Inference works, and its OpenAI-compatible API.
-
Capabilities
Use tool calling, structured output, and reasoning.
-
Reference
Look up exact endpoints, parameters, and pricing.
-
Platform
Manage API keys, billing, credits, and usage.
-
Resources
Release notes and answers to common questions.