Skip to content

Reasoning

Some models produce explicit reasoning (a chain of thought) before their final answer. This guide covers how to use reasoning models on Mango Inference and how to read the reasoning output.

When to use it

Reasoning models trade latency and token cost for stronger performance on multi-step problems: math, code, planning. For simple tasks, a standard model is faster and cheaper.

Make a reasoning request

resp = client.chat.completions.create(
    model="moonshotai/Kimi-K3",
    messages=[{"role": "user", "content": "If a train travels 60km in 45 minutes, what is its speed in km/h?"}],
)

Note what is not in that request: nothing enables reasoning. On a reasoning-capable model it happens by default. There is no reasoning_effort-style parameter documented for Mango Inference, so you get the model's own behaviour. If you need shorter answers, cap max_tokens or choose a non-reasoning model; you cannot currently dial the reasoning down.

Reading the output

Reasoning and the answer come back in the same message, in separate fields:

Field Contains
choices[].message.reasoning_content The model's reasoning. Not an OpenAI field.
choices[].message.content The final answer.
msg = resp.choices[0].message
print(msg.reasoning_content)  # how it got there
print(msg.content)            # what to show the user

Show content to users. Reasoning is working notes: it can contradict the final answer, wander, or restate the prompt, and it is not written to be read.

OpenAI SDKs may not expose reasoning_content

It is not part of the OpenAI schema. Depending on your SDK version it may be dropped from the typed object even though it was on the wire. Read it from the raw response (for example resp.model_dump()) if the attribute is missing. See OpenAI compatibility.

What reasoning costs

Reasoning is not free, and it is not separately priced. It is generated tokens, reported separately and counted in the same total:

"usage": {
  "prompt_tokens": 14,
  "completion_tokens": 247,
  "reasoning_tokens": 237,
  "total_tokens": 261
}

In that response (a one-line greeting), 237 of the 247 generated tokens were reasoning. reasoning_tokens is a breakdown of completion_tokens, not an addition to it, so total_tokens is still prompt + completion. Budget accordingly: a reasoning model can cost an order of magnitude more than a standard one for the same visible output, which is the real reason to reserve it for problems that need it.

Supported models

Reasoning is a property of the model. List the models your key can call with GET /v1/models (see Models and pricing) and check the model card, or just make a call and look for reasoning_content in the response, which is the definitive test.