Messages¶
POST /v1/messages
Generate a model response for a conversation using the Anthropic Messages API request/response shape. This lets Anthropic-ecosystem tools (Claude Code, agent frameworks built on the Anthropic SDK, etc.) point at Mango Inference without changing their request format. See Coding agent setup for tool-specific setup guides.
Request¶
POST https://api.mangoboost.io/v1/messages
x-api-key: <your-api-key>
Content-Type: application/json
The API key is the same Mango Inference API key used for the OpenAI-compatible endpoint. Only the header name differs. See Authentication.
Anthropic's own API requires an anthropic-version header, and every Anthropic
SDK sets it for you, so tools like Claude Code need no extra configuration.
Mango Inference's behaviour when the header is absent is not documented. If you
are calling this endpoint from a hand-rolled HTTP client, send
anthropic-version: 2023-06-01 rather than relying on it being optional.
Body parameters¶
The request body follows the Anthropic Messages schema. The parameters below are the ones these docs cover; a Messages-shaped field not listed here is passed through to the model server rather than validated by Mango Inference, so treat anything outside this table as unsupported until you have tested it.
| Parameter | Type | Required | Description |
|---|---|---|---|
model |
string | Yes | Model ID. See Models and pricing. |
messages |
array | Yes | Conversation messages (role + content). |
max_tokens |
integer | Yes | Maximum tokens to generate. |
system |
string | No | System prompt. |
temperature |
number | No | Sampling temperature. |
top_p |
number | No | Nucleus sampling cutoff. |
stream |
boolean | No | Stream the response as server-sent events. Defaults to false. |
stop_sequences |
array | No | Stop sequence(s). |
tools |
array | No | Tool definitions. See Tool calling. |
tool_choice |
object | No | How the model selects tools. |
As with Chat Completions, sampling parameters you omit fall back to the serving defaults for the model, which are not published. Set the ones your application depends on explicitly.
Response¶
Responses use the Anthropic Messages response shape:
{
"id": "msg_01XFDUDYJgAACzvnptvVoYEL",
"type": "message",
"role": "assistant",
"model": "zai-org/GLM-5.3",
"content": [
{"type": "text", "text": "Hello! How can I help you today?"}
],
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {"input_tokens": 14, "output_tokens": 10}
}
Response fields¶
| Field | Description |
|---|---|
content[].text |
The generated text. |
stop_reason |
Why generation stopped (end_turn, max_tokens, tool_use, …). |
usage |
Token counts for the request. |
Errors¶
Authentication failures return HTTP 401 with the Anthropic error envelope,
not the OpenAI one:
{"type": "error", "error": {"type": "authentication_error", "message": "invalid API key"}}
A client that talks to both endpoint families has to handle both shapes. See Reliability and error handling.