Overview¶
Mango Inference is MangoBoost's platform for running large language models as a service. You send requests to an OpenAI-compatible API, and Mango Inference handles model serving, scaling, and reliability. There is no infrastructure to manage.
This page orients you before the Quickstart.
What you can do¶
- Call open and hosted models through a single, OpenAI-compatible endpoint.
- Use capabilities such as tool calling, structured output, and reasoning.
- Plug Mango Inference into existing tooling via the OpenAI SDK and coding agent setup.
- Run a coding agent (Claude Code, OpenCode, OpenHands) against your own repository. See Coding agent setup.
How it works¶
Your client sends an HTTP request to https://api.mangoboost.io/v1, naming a model
and carrying your API key. Mango Inference authenticates the key, checks that it
may call that model, and routes the request to shared capacity already running
those weights. The model generates a response, and you get it back, either
in one reply or streamed token by token if you asked for that. Each response
reports what it consumed, and those tokens are drawn from your balance.
There is no step in that path you operate: no cold start to wait out, no capacity to reserve beforehand, nothing running between requests.
Mango Inference exposes serverless Model APIs: you pick a model, send a request, and pay for what you use. See Platform concepts for the core objects and terms.
Next steps¶
-
Quickstart
Make your first API call in a few minutes.
-
Coding agent setup
Point Claude Code, OpenCode, or OpenHands at Mango Inference.
-
Platform concepts
Learn the mental model behind Mango Inference.
-
API reference
Look up exact endpoints and parameters.