Skip to content

Chat Completions

POST /v1/chat/completions

Generate a model response for a conversation. This endpoint is OpenAI-compatible.

Request

POST https://api.mangoboost.io/v1/chat/completions
Authorization: Bearer <your-api-key>
Content-Type: application/json

Body parameters

The request body follows the OpenAI Chat Completions schema. The parameters below are the ones these docs cover; an OpenAI-shaped field not listed here is passed through to the model server rather than validated by Mango Inference, so treat anything outside this table as unsupported until you have tested it.

Parameter Type Required Description
model string Yes Model ID. See Models and pricing.
messages array Yes Conversation messages (role + content).
max_tokens integer No Maximum tokens to generate.
temperature number No Sampling temperature.
top_p number No Nucleus sampling cutoff.
stream boolean No Stream tokens as server-sent events. Defaults to false.
stop string / array No Stop sequence(s).
tools array No Tool definitions. See Tool calling.
tool_choice string / object No How the model selects tools.
response_format object No Constrain output. See Structured output.

Sampling parameters you omit fall back to the serving defaults for the model. Those defaults are not published and can differ per model, so set the ones your application depends on explicitly rather than relying on them.

Response

{
  "id": "e9c0b646861d4c0085503fe54d796afa",
  "object": "chat.completion",
  "created": 1785742283,
  "model": "zai-org/GLM-5.3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! How can I help you today?",
        "reasoning_content": "…",
        "tool_calls": null
      },
      "logprobs": null,
      "finish_reason": "stop",
      "matched_stop": 154827
    }
  ],
  "usage": {
    "prompt_tokens": 14,
    "total_tokens": 261,
    "completion_tokens": 247,
    "prompt_tokens_details": null,
    "reasoning_tokens": 237
  },
  "metadata": {
    "weight_version": "default"
  }
}

Response fields

Field Description
choices[].message.content The generated text.
choices[].message.reasoning_content Reasoning the model produced before its answer, on reasoning-capable models. Not an OpenAI field. See Reasoning.
choices[].message.tool_calls Tool calls the model wants you to run, or null. See Tool calling.
choices[].finish_reason Why generation stopped (stop, length, tool_calls, …).
choices[].matched_stop Which stop condition ended generation: a token ID, or the stop string that matched. Not an OpenAI field.
usage Token counts for the request. reasoning_tokens counts reasoning separately, and is included in completion_tokens.
metadata.weight_version The model weights that served the request. Not an OpenAI field.

The three fields marked "not an OpenAI field" are additions. OpenAI SDKs ignore unknown response fields, so they do not break existing client code. However, code that round-trips a response through a strict schema may need to allow them.

Errors

See Reliability and error handling.