Skip to content

Messages

POST /v1/messages

Generate a model response for a conversation using the Anthropic Messages API request/response shape. This lets Anthropic-ecosystem tools (Claude Code, agent frameworks built on the Anthropic SDK, etc.) point at Mango Inference without changing their request format. See Coding agent setup for tool-specific setup guides.

Request

POST https://api.mangoboost.io/v1/messages
x-api-key: <your-api-key>
Content-Type: application/json

The API key is the same Mango Inference API key used for the OpenAI-compatible endpoint. Only the header name differs. See Authentication.

Anthropic's own API requires an anthropic-version header, and every Anthropic SDK sets it for you, so tools like Claude Code need no extra configuration. Mango Inference's behaviour when the header is absent is not documented. If you are calling this endpoint from a hand-rolled HTTP client, send anthropic-version: 2023-06-01 rather than relying on it being optional.

Body parameters

The request body follows the Anthropic Messages schema. The parameters below are the ones these docs cover; a Messages-shaped field not listed here is passed through to the model server rather than validated by Mango Inference, so treat anything outside this table as unsupported until you have tested it.

Parameter Type Required Description
model string Yes Model ID. See Models and pricing.
messages array Yes Conversation messages (role + content).
max_tokens integer Yes Maximum tokens to generate.
system string No System prompt.
temperature number No Sampling temperature.
top_p number No Nucleus sampling cutoff.
stream boolean No Stream the response as server-sent events. Defaults to false.
stop_sequences array No Stop sequence(s).
tools array No Tool definitions. See Tool calling.
tool_choice object No How the model selects tools.

As with Chat Completions, sampling parameters you omit fall back to the serving defaults for the model, which are not published. Set the ones your application depends on explicitly.

Response

Responses use the Anthropic Messages response shape:

{
  "id": "msg_01XFDUDYJgAACzvnptvVoYEL",
  "type": "message",
  "role": "assistant",
  "model": "zai-org/GLM-5.3",
  "content": [
    {"type": "text", "text": "Hello! How can I help you today?"}
  ],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {"input_tokens": 14, "output_tokens": 10}
}

Response fields

Field Description
content[].text The generated text.
stop_reason Why generation stopped (end_turn, max_tokens, tool_use, …).
usage Token counts for the request.

Errors

Authentication failures return HTTP 401 with the Anthropic error envelope, not the OpenAI one:

{"type": "error", "error": {"type": "authentication_error", "message": "invalid API key"}}

A client that talks to both endpoint families has to handle both shapes. See Reliability and error handling.