Skip to content

FAQ

Frequently asked questions about Mango Inference.

About the product

What is Mango Inference?

MangoBoost's platform for running large language models as a service. You send requests to an OpenAI-compatible API and Mango Inference handles model serving, scaling, and reliability. See Overview.

Do I have to manage any servers or GPUs?

No. Model APIs are serverless: you call a model by ID, Mango Inference handles capacity, and you pay per use. There are no replicas or scaling to configure.

Where should I start?

The Quickstart: first API call, then choosing a model, streaming, and controlling the output. After that, Coding agent setup if you want an agent working in your repository.

API and compatibility

Is the API OpenAI-compatible?

Yes. Point an existing OpenAI client at Mango Inference by changing two things: the base URL and the API key. See OpenAI compatibility.

What is the base URL?

https://api.mangoboost.io/v1

Which endpoints are supported?

Four: POST /v1/chat/completions (the primary surface), POST /v1/messages for Anthropic-shaped clients, POST /v1/messages/count_tokens for token counting, and GET /v1/models. See the API reference.

Which OpenAI endpoints are not supported?

Everything except chat. /v1/embeddings, /v1/completions (legacy), /v1/responses, /v1/moderations, images, audio, files, and batches all return 404. The full boundary is in OpenAI compatibility.

The one that catches people out is embeddings: a RAG framework configured with a single OpenAI-compatible provider will route both chat and embeddings to Mango Inference, and the embedding calls fail. Point them elsewhere.

Does Mango Inference support the Batch API?

No. The OpenAI Batch API (/v1/batches, and /v1/files for batch processing) is not supported; batch-related paths return 404. If you need to process large volumes of requests asynchronously, submit them as individual /v1/chat/completions requests with your own concurrency control. See Rate limits for the ceilings that apply.

Can I use the official OpenAI SDKs?

Yes, with the base URL and key overridden. See the OpenAI SDK guide. Support is partial at MVP: Chat Completions is the primary supported surface, and other SDK features may not be available yet.

Does Mango Inference work with LangChain, LlamaIndex, or other frameworks?

Any tool with an "OpenAI-compatible" or custom base URL setting can target Mango Inference. See OpenAI compatibility for the supported endpoints and known differences.

Four tools have a written setup guide: the OpenAI SDK, Claude Code, OpenCode, and OpenHands. LangChain, LlamaIndex, the Vercel AI SDK and LiteLLM should work through the same base-URL pattern but have not been verified here. Check whether the framework lets you override the base URL and whether it needs endpoints beyond chat (see OpenAI compatibility).

Do I need to change my request or response parsing code?

Requests and responses follow the OpenAI schema, so existing code works with minimal changes.

The differences are additions, not changes: responses may carry reasoning_content, matched_stop, and metadata.weight_version, and reasoning models report usage.reasoning_tokens. SDKs ignore unknown fields, so normal client code is unaffected. Strict schema validation on your side is the thing that will complain. Model IDs are the other change: they are upstream repository IDs like zai-org/GLM-5.3, not OpenAI model names. Full list: Known differences.

Can I stream responses?

Yes. Set stream: true to receive tokens as server-sent events as they are generated. See Quickstart.

Does Mango Inference support tool calling, structured output, and reasoning?

Yes, through the OpenAI-compatible tools, response_format, and reasoning-capable models:

Support varies by model. See Models and pricing.

Models

Which models are available?

Models and pricing is the authoritative list. Pass a model's ID in the model field of your request.

How do I request a model that isn't available?

Ask in Discord and tell us which model you want deployed. We take requests for new and popular models. See Platform support.

Which model should I use?

Browse Models and pricing and compare context length, capabilities, and price. Reasoning models trade latency and token cost for stronger multi-step performance; a standard model is faster and cheaper for simple tasks. See Reasoning.

API keys

How do I authenticate?

Send your API key as a bearer token on every request:

Authorization: Bearer <your-api-key>

I lost my API key. Can I see it again?

No. A key is shown only once when created. Create a new key and update your applications. See API keys.

How should I store my key?

In an environment variable or a secrets manager. Never commit a key to source control.

Can I scope a key to one model or make it read-only?

No. Keys are unscoped: a key either works for the account's inference or it does not. Isolate environments by issuing separate keys, not by scoping one.

How do I rotate or revoke a key?

Both are done in the console. Revoking is deleting the key; requests using it fail with 401 immediately after. To rotate without downtime, create the new key first, deploy it, confirm traffic has moved, then delete the old one. There is no built-in grace period. See API keys.

Billing and usage

How am I charged?

Usage draws down your credits/balance at the rates in Models and pricing.

Standard accounts are prepaid: you hold a balance and per-token usage meters against it, so an account can never spend more than what is on it. Every response reports its own token counts in usage. Reasoning tokens are ordinary generated tokens: reasoning_tokens is a breakdown of completion_tokens, not an extra charge on top. See Billing and payments.

What happens if I run out of credits?

Requests may be rejected, for example with an HTTP 402. Top up your balance; see also Account suspension.

How do I see what I'm spending?

In the console. See Usage and cost breakdown.

Which payment methods are supported?

See Billing and payments.

The accepted methods are not published here. Set payment up (or ask what is possible, including invoicing for larger accounts) through Discord or support@mangoboost.io before you depend on a specific method.

Limits and errors

What are the rate limits?

See Rate limits for the enforced units, scope, and defaults.

No numeric default is published, and limits are not identical for every account. Rather than hard-code a figure, handle 429 with exponential backoff and jitter and cap your own concurrency. Once you do, the exact ceiling stops mattering. To find out what applies to your key, or to ask for more, contact Discord or support@mangoboost.io; include your models and your expected sustained and peak rates. See Rate limits.

I'm getting HTTP 429. What should I do?

Back off and retry with exponential backoff and jitter, and respect the Retry-After header if one is present. See Reliability and error handling.

Which errors are safe to retry?

Only transient failures: 429 and 5xx. A 400 means the request itself is wrong; fix it rather than retrying unchanged. A 401 means the API key is missing or bad. See Reliability and error handling.

My long generation times out. What can I do?

Set a client-side timeout appropriate to your max_tokens, and prefer streaming so partial output arrives incrementally.

Data and security

Are my prompts used to train models?

See Data privacy and security.

Not answered yet. These docs will not guess, because a retention or training claim is a commitment. Mango Inference has not published its terms for prompt and completion retention or for training use, and has published no compliance certifications.

What is settled is the mechanical side: the API is HTTPS-only (TLS 1.3, with plain HTTP redirected), every request is authenticated, and a key only reaches the models it is entitled to.

If your use depends on a specific answer (a data-processing agreement, a zero-retention arrangement, a deletion SLA), get it in writing from support@mangoboost.io before building on the platform. See Data privacy and security.

Which regions do deployments run in?

Not yet published. See Platform support. If you have a specific region requirement, tell us through Discord or support@mangoboost.io.

Getting help

Where do I ask a question or report a bug?

Join the Mango Inference Discord community, or email support@mangoboost.io. See Platform support.

How do I request a feature?

Ask in Discord. Feature requests, model requests, and bug reports all go through the same channel.