Reliability and error handling¶
How to build clients that handle errors gracefully and keep working under load.
Error responses¶
Mango Inference returns standard HTTP status codes with a JSON body. The body's shape depends on which endpoint family you called, so a client that uses both has to parse both.
/v1/models, /v1/chat/completions
{"error": {"message": "invalid API key", "type": "invalid_request_error"}}
/v1/messages
{"type": "error", "error": {"type": "authentication_error", "message": "invalid API key"}}
Read the status code first and the body second: the status is what your retry
logic should branch on, and the message is for your logs.
Status codes¶
| Status | Meaning | What to do |
|---|---|---|
400 |
Invalid request | Fix the request; do not retry unchanged. |
401 |
Missing, malformed, or unknown API key | Check your API key. Not retryable. |
402 |
API key has insufficient balance | Top up your balance. |
403 |
Key disabled, expired, under legal hold, or model not allowed | Check the error message for the specific cause. Not retryable. |
404 |
No such endpoint | You called a path Mango Inference does not serve. See OpenAI compatibility. |
405 |
Wrong method for the path | /v1/chat/completions and /v1/messages are POST; /v1/models is GET. |
429 |
Rate limited | Back off and retry. See rate limits. |
500 |
Internal server error | Retry with backoff. |
502 |
Upstream unreachable | The middleware could not reach any inference engine. Retry with backoff. |
503 |
Service unavailable (capacity, maintenance, or backend) | Retry with backoff. Respect Retry-After if present. |
529 |
Upstream overloaded | The inference engine is shedding load. Retry with backoff. Respect Retry-After. |
A 404 here means "not a route", not "model not found". An unknown model ID
comes back as a 4xx from the endpoint itself, with an error body. If you get
plain-text 404 page not found rather than JSON, you have the URL wrong; the
usual cause is a /v1 that is doubled or missing.
Error types and codes¶
Every error response carries a JSON body with a type field, and some also
carry a machine-readable code that names the specific failure. A client that
branches on the code can distinguish errors that share an HTTP status. For
example, it can tell a 429 from the platform's own rate limiter
(rate_limit_error) apart from a 429 relayed from an upstream engine
(upstream_error with code upstream_rate_limited).
The OpenAI-compatible envelope includes code when the middleware sets one; the
Anthropic envelope derives its error type from the HTTP status and does not
carry a separate code.
Authentication and key errors¶
| Status | Error type | Error code | Message |
|---|---|---|---|
401 |
invalid_request_error |
(none) | invalid API key |
402 |
invalid_request_error |
(none) | API key has insufficient balance |
403 |
invalid_request_error |
(none) | API key is under legal hold; an org admin must accept the pending document in the dashboard |
403 |
invalid_request_error |
(none) | API key is disabled |
403 |
invalid_request_error |
(none) | API key has expired |
503 |
internal_error |
(none) | authentication backend unavailable |
Model authorization errors¶
| Status | Error type | Error code | Message |
|---|---|---|---|
403 |
invalid_request_error |
(none) | model {model} is not allowed for this API key |
403 |
invalid_request_error |
(none) | model {model} is not available to this organization |
404 |
invalid_request_error |
(none) | model {model} is not available on this API |
503 |
internal_error |
(none) | no upstream host configured for model |
503 |
invalid_request_error |
model_under_maintenance |
model {model} is under maintenance |
503 |
internal_error |
(none) | model entitlements are temporarily unavailable |
Rate limiting¶
| Status | Error type | Error code | Message |
|---|---|---|---|
429 |
rate_limit_error |
(names the limit dimension) | rate limit exceeded: {N} {resource} per minute for this API key |
A 429 from the platform's own rate limiter carries these headers:
| Header | Meaning |
|---|---|
Retry-After |
Seconds until the limit resets. |
X-RateLimit-Limit |
The limit on the dimension that tripped. |
X-RateLimit-Remaining |
Remaining requests in the window. |
X-RateLimit-Reset |
Seconds until reset (same value as Retry-After). |
X-RateLimit-Resource |
Which limit dimension was hit (e.g. requests, tokens). |
Upstream and capacity errors¶
| Status | Error type | Error code | Message |
|---|---|---|---|
429 |
upstream_error |
upstream_rate_limited |
upstream inference engine is rate limiting requests |
500 |
internal_error |
(none) | upstream inference engine returned an error |
502 |
upstream_error |
(none) | failed to reach upstream inference engine |
503 |
upstream_error |
upstream_at_capacity |
model {model} is at capacity |
503 |
upstream_error |
upstream_draining |
model {model} is not currently accepting requests |
503 |
upstream_error |
(none) | model {model} is temporarily unavailable |
529 |
upstream_error |
upstream_overloaded |
upstream inference engine is overloaded |
A 429 or 529 relayed from an upstream engine carries a Retry-After header
if the engine provided one. The 503 capacity and availability errors also
carry Retry-After.
Upstream 400 relay¶
When the upstream engine rejects the request with a 400, the middleware
relays the engine's JSON error body verbatim if it is valid JSON; otherwise it
wraps the message in its own envelope. The error code upstream_invalid_request
distinguishes a request the engine rejected from one the middleware's own
validator rejected.
In-stream errors¶
A streaming response that started with 200 can still fail mid-stream: the
engine sends an error frame inside the SSE stream after tokens have already been
generated. The client sees the stream end; the error is captured in the platform's
logs and usage records. Tokens generated before the error frame are billed.
Retries and backoff¶
- Retry only transient failures:
429and5xx. Never retry a400or401unchanged; the second attempt fails identically and costs you a round trip. - Use exponential backoff with jitter. Without jitter, a fleet that hits a limit together retries together and hits it again.
- Respect a
Retry-Afterheader if one is present, in preference to your own schedule. - Cap total attempts and total elapsed time. An unbounded retry loop against a rate limit is indistinguishable from an attack on your own quota.
- Streamed responses can fail mid-stream, after a
200. Decide whether a partial completion is usable before you retry, because retrying charges you for the tokens twice.
Timeouts and streaming¶
- Set a client-side timeout appropriate to your
max_tokens. Reasoning models generate far more tokens than their visible answer suggests, so a timeout tuned on a standard model will fire early on one. See Reasoning. - For long generations, prefer streaming so partial output arrives incrementally and an idle connection is distinguishable from a slow one.
Reporting a problem¶
Every response carries x-request-id and x-trace-id. Log them. When
something looks like a platform fault rather than a client bug, quote them,
with the timestamp and the model ID, in Discord or to
support@mangoboost.io. They are how the team finds
your specific request. See Platform support.