Skip to content

Platform concepts

The mental model behind Mango Inference: the objects you work with and the terms used throughout these docs.

Key concepts

Concept What it is
Model A hosted LLM you call by its model ID.
Model API The endpoint you send requests to: OpenAI-compatible (/v1/chat/completions) or Anthropic-compatible (/v1/messages). See Model APIs.
API key The credential that authenticates your requests. See API keys.
Account Your organization's workspace, holding keys, usage, and billing.
Credits / balance Prepaid funds consumed by usage. See Credits and balance.
Rate limit The ceiling on requests/tokens per unit time. See Rate limits.

How it fits together

An account is the top-level container: it holds your keys, your balance, and your usage history. Inside it you create one or more API keys.

A key is what a request carries, sent as Authorization: Bearer … for the OpenAI-compatible endpoints, or as x-api-key for the Anthropic-compatible one (Authentication). The key identifies both who is calling and which models they may call, so GET /v1/models returns the catalogue for that key, not a global list.

Each accepted request runs on shared serverless capacity, subject to the account's rate limits, and reports what it consumed in the response's usage object. Those tokens are priced and drawn down from the account's balance. Run the balance to zero and requests stop being served. See Credits and balance and Account suspension.

Nothing in that chain involves capacity you provision or pay for while idle: between requests, an account costs nothing.