Skip to content

Overview

Mango Inference is MangoBoost's platform for running large language models as a service. You send requests to an OpenAI-compatible API, and Mango Inference handles model serving, scaling, and reliability. There is no infrastructure to manage.

This page orients you before the Quickstart.

What you can do

How it works

Your client sends an HTTP request to https://api.mangoboost.io/v1, naming a model and carrying your API key. Mango Inference authenticates the key, checks that it may call that model, and routes the request to shared capacity already running those weights. The model generates a response, and you get it back, either in one reply or streamed token by token if you asked for that. Each response reports what it consumed, and those tokens are drawn from your balance.

There is no step in that path you operate: no cold start to wait out, no capacity to reserve beforehand, nothing running between requests.

Mango Inference exposes serverless Model APIs: you pick a model, send a request, and pay for what you use. See Platform concepts for the core objects and terms.

Next steps

  • Quickstart


    Make your first API call in a few minutes.

    Quickstart

  • Coding agent setup


    Point Claude Code, OpenCode, or OpenHands at Mango Inference.

    Coding agents

  • Platform concepts


    Learn the mental model behind Mango Inference.

    Concepts

  • API reference


    Look up exact endpoints and parameters.

    Reference