Skip to content

Reliability and error handling

How to build clients that handle errors gracefully and keep working under load.

Error responses

Mango Inference returns standard HTTP status codes with a JSON body. The body's shape depends on which endpoint family you called, so a client that uses both has to parse both.

/v1/models, /v1/chat/completions

{"error": {"message": "invalid API key", "type": "invalid_request_error"}}

/v1/messages

{"type": "error", "error": {"type": "authentication_error", "message": "invalid API key"}}

Read the status code first and the body second: the status is what your retry logic should branch on, and the message is for your logs.

Status codes

Status Meaning What to do
400 Invalid request Fix the request; do not retry unchanged.
401 Missing, malformed, or unknown API key Check your API key. Not retryable.
402 API key has insufficient balance Top up your balance.
403 Key disabled, expired, under legal hold, or model not allowed Check the error message for the specific cause. Not retryable.
404 No such endpoint You called a path Mango Inference does not serve. See OpenAI compatibility.
405 Wrong method for the path /v1/chat/completions and /v1/messages are POST; /v1/models is GET.
429 Rate limited Back off and retry. See rate limits.
500 Internal server error Retry with backoff.
502 Upstream unreachable The middleware could not reach any inference engine. Retry with backoff.
503 Service unavailable (capacity, maintenance, or backend) Retry with backoff. Respect Retry-After if present.
529 Upstream overloaded The inference engine is shedding load. Retry with backoff. Respect Retry-After.

A 404 here means "not a route", not "model not found". An unknown model ID comes back as a 4xx from the endpoint itself, with an error body. If you get plain-text 404 page not found rather than JSON, you have the URL wrong; the usual cause is a /v1 that is doubled or missing.

Error types and codes

Every error response carries a JSON body with a type field, and some also carry a machine-readable code that names the specific failure. A client that branches on the code can distinguish errors that share an HTTP status. For example, it can tell a 429 from the platform's own rate limiter (rate_limit_error) apart from a 429 relayed from an upstream engine (upstream_error with code upstream_rate_limited).

The OpenAI-compatible envelope includes code when the middleware sets one; the Anthropic envelope derives its error type from the HTTP status and does not carry a separate code.

Authentication and key errors

Status Error type Error code Message
401 invalid_request_error (none) invalid API key
402 invalid_request_error (none) API key has insufficient balance
403 invalid_request_error (none) API key is under legal hold; an org admin must accept the pending document in the dashboard
403 invalid_request_error (none) API key is disabled
403 invalid_request_error (none) API key has expired
503 internal_error (none) authentication backend unavailable

Model authorization errors

Status Error type Error code Message
403 invalid_request_error (none) model {model} is not allowed for this API key
403 invalid_request_error (none) model {model} is not available to this organization
404 invalid_request_error (none) model {model} is not available on this API
503 internal_error (none) no upstream host configured for model
503 invalid_request_error model_under_maintenance model {model} is under maintenance
503 internal_error (none) model entitlements are temporarily unavailable

Rate limiting

Status Error type Error code Message
429 rate_limit_error (names the limit dimension) rate limit exceeded: {N} {resource} per minute for this API key

A 429 from the platform's own rate limiter carries these headers:

Header Meaning
Retry-After Seconds until the limit resets.
X-RateLimit-Limit The limit on the dimension that tripped.
X-RateLimit-Remaining Remaining requests in the window.
X-RateLimit-Reset Seconds until reset (same value as Retry-After).
X-RateLimit-Resource Which limit dimension was hit (e.g. requests, tokens).

Upstream and capacity errors

Status Error type Error code Message
429 upstream_error upstream_rate_limited upstream inference engine is rate limiting requests
500 internal_error (none) upstream inference engine returned an error
502 upstream_error (none) failed to reach upstream inference engine
503 upstream_error upstream_at_capacity model {model} is at capacity
503 upstream_error upstream_draining model {model} is not currently accepting requests
503 upstream_error (none) model {model} is temporarily unavailable
529 upstream_error upstream_overloaded upstream inference engine is overloaded

A 429 or 529 relayed from an upstream engine carries a Retry-After header if the engine provided one. The 503 capacity and availability errors also carry Retry-After.

Upstream 400 relay

When the upstream engine rejects the request with a 400, the middleware relays the engine's JSON error body verbatim if it is valid JSON; otherwise it wraps the message in its own envelope. The error code upstream_invalid_request distinguishes a request the engine rejected from one the middleware's own validator rejected.

In-stream errors

A streaming response that started with 200 can still fail mid-stream: the engine sends an error frame inside the SSE stream after tokens have already been generated. The client sees the stream end; the error is captured in the platform's logs and usage records. Tokens generated before the error frame are billed.

Retries and backoff

  • Retry only transient failures: 429 and 5xx. Never retry a 400 or 401 unchanged; the second attempt fails identically and costs you a round trip.
  • Use exponential backoff with jitter. Without jitter, a fleet that hits a limit together retries together and hits it again.
  • Respect a Retry-After header if one is present, in preference to your own schedule.
  • Cap total attempts and total elapsed time. An unbounded retry loop against a rate limit is indistinguishable from an attack on your own quota.
  • Streamed responses can fail mid-stream, after a 200. Decide whether a partial completion is usable before you retry, because retrying charges you for the tokens twice.

Timeouts and streaming

  • Set a client-side timeout appropriate to your max_tokens. Reasoning models generate far more tokens than their visible answer suggests, so a timeout tuned on a standard model will fire early on one. See Reasoning.
  • For long generations, prefer streaming so partial output arrives incrementally and an idle connection is distinguishable from a slow one.

Reporting a problem

Every response carries x-request-id and x-trace-id. Log them. When something looks like a platform fault rather than a client bug, quote them, with the timestamp and the model ID, in Discord or to support@mangoboost.io. They are how the team finds your specific request. See Platform support.