Platform concepts¶
The mental model behind Mango Inference: the objects you work with and the terms used throughout these docs.
Key concepts¶
| Concept | What it is |
|---|---|
| Model | A hosted LLM you call by its model ID. |
| Model API | The endpoint you send requests to: OpenAI-compatible (/v1/chat/completions) or Anthropic-compatible (/v1/messages). See Model APIs. |
| API key | The credential that authenticates your requests. See API keys. |
| Account | Your organization's workspace, holding keys, usage, and billing. |
| Credits / balance | Prepaid funds consumed by usage. See Credits and balance. |
| Rate limit | The ceiling on requests/tokens per unit time. See Rate limits. |
How it fits together¶
An account is the top-level container: it holds your keys, your balance, and your usage history. Inside it you create one or more API keys.
A key is what a request carries, sent as Authorization: Bearer … for the
OpenAI-compatible endpoints, or as x-api-key for the Anthropic-compatible one
(Authentication). The key identifies both
who is calling and which models they may call, so GET /v1/models returns
the catalogue for that key, not a global list.
Each accepted request runs on shared serverless capacity, subject to the
account's rate limits, and reports what it consumed in the response's
usage object. Those tokens are priced and drawn down from the account's
balance. Run the balance to zero and requests stop being served. See
Credits and balance and
Account suspension.
Nothing in that chain involves capacity you provision or pay for while idle: between requests, an account costs nothing.