Skip to content

Quickstart

Make your first request to the Mango Inference API, then pick a model, stream the response, and control the output, all in one sitting. For the platform view of serverless Model APIs, see Model APIs.

Prerequisites

  • A Mango Inference account.
  • An API key.
  • curl, or Python 3.8+ with the openai package for the SDK example.

1. Set your API key

export MANGOINFERENCE_API_KEY="<your-api-key>"

2. Send your first request

Mango Inference is OpenAI-compatible, so you call it the same way you call the OpenAI API. Only the base URL and key change.

curl https://api.mangoboost.io/v1/chat/completions \
  -H "Authorization: Bearer $MANGOINFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zai-org/GLM-5.3",
    "messages": [
      {"role": "user", "content": "Hello, Mango Inference!"}
    ]
  }'
from openai import OpenAI

client = OpenAI(
    base_url="https://api.mangoboost.io/v1",
    api_key="<your-api-key>",  # or set OPENAI_API_KEY
)

resp = client.chat.completions.create(
    model="zai-org/GLM-5.3",
    messages=[{"role": "user", "content": "Hello, Mango Inference!"}],
)
print(resp.choices[0].message.content)

3. Read the response

{
  "id": "e9c0b646861d4c0085503fe54d796afa",
  "object": "chat.completion",
  "created": 1785742283,
  "model": "zai-org/GLM-5.3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! How can I help you today?",
        "reasoning_content": "1.  **Analyze the Input**:\n    *   Input: \"Hello?\"\n    *   Tone/Intent: Greeting, checking for connection/presence, possibly inquisitive.\n    *   Expected Output: A friendly greeting, acknowledging presence, and an offer to help.\n\n2.  **Determine the Persona**:\n    *   AI assistant.\n    *   Helpful, polite, responsive.\n\n3.  **Draft Responses**:\n    *   *Option 1*: Hello! How can I help you today?\n    *   *Option 2*: Hi there! I'm here. What can I do for you?\n    *   *Option 3*: Hello! I'm listening. How can I assist you today?\n\n4.  **Select the Best Response**:\n    *   Option 3 is good, but Option 1 is classic and effective. Let's go with a warm, friendly variation: \"Hello! How can I help you today?\" or \"Hi there! How can I assist you today?\"\n\n5.  **Final Polish**: \"Hello! How can I help you today?\" (Simple, polite, direct).",
        "tool_calls": null
      },
      "logprobs": null,
      "finish_reason": "stop",
      "matched_stop": 154827
    }
  ],
  "usage": {
    "prompt_tokens": 14,
    "total_tokens": 261,
    "completion_tokens": 247,
    "prompt_tokens_details": null,
    "reasoning_tokens": 237
  },
  "metadata": {
    "weight_version": "default"
  }
}

choices[].message.content is the answer. The rest is worth a glance now so it does not surprise you later: reasoning_content appears on reasoning-capable models and is working notes rather than output, and usage.reasoning_tokens is a breakdown of completion_tokens, not an extra charge on top. Full field list: Chat Completions.

4. Choose a model

Ask the API which models your key can call, then pass one of the returned id values in the model field:

curl https://api.mangoboost.io/v1/models \
  -H "Authorization: Bearer $MANGOINFERENCE_API_KEY"

The catalogue changes as models are added, so read it rather than hard-coding a list. See Models and pricing. The examples here use zai-org/GLM-5.3.

5. Stream the response

Set stream: true to receive tokens as they are generated.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.mangoboost.io/v1",
    api_key="<your-api-key>",
)

stream = client.chat.completions.create(
    model="zai-org/GLM-5.3",
    messages=[{"role": "user", "content": "Write a haiku about mangoes."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

6. Control the output

Common parameters (see the Chat Completions reference for the full list):

Parameter Purpose
temperature Randomness of the output.
max_tokens Upper bound on generated tokens.
top_p Nucleus sampling cutoff.
stop Stop sequences.

Anything you leave unset falls back to the serving default for that model, and those defaults are not published. Set the parameters your application depends on explicitly rather than assuming OpenAI's values apply.

Next steps