Reasoning¶
Some models produce explicit reasoning (a chain of thought) before their final answer. This guide covers how to use reasoning models on Mango Inference and how to read the reasoning output.
When to use it¶
Reasoning models trade latency and token cost for stronger performance on multi-step problems: math, code, planning. For simple tasks, a standard model is faster and cheaper.
Make a reasoning request¶
resp = client.chat.completions.create(
model="moonshotai/Kimi-K3",
messages=[{"role": "user", "content": "If a train travels 60km in 45 minutes, what is its speed in km/h?"}],
)
Note what is not in that request: nothing enables reasoning. On a
reasoning-capable model it happens by default. There is no
reasoning_effort-style parameter documented for Mango Inference, so you get the
model's own behaviour. If you need shorter answers, cap max_tokens or choose a
non-reasoning model; you cannot currently dial the reasoning down.
Reading the output¶
Reasoning and the answer come back in the same message, in separate fields:
| Field | Contains |
|---|---|
choices[].message.reasoning_content |
The model's reasoning. Not an OpenAI field. |
choices[].message.content |
The final answer. |
msg = resp.choices[0].message
print(msg.reasoning_content) # how it got there
print(msg.content) # what to show the user
Show content to users. Reasoning is working notes: it can contradict the
final answer, wander, or restate the prompt, and it is not written to be read.
OpenAI SDKs may not expose reasoning_content
It is not part of the OpenAI schema. Depending on your SDK version it may
be dropped from the typed object even though it was on the wire. Read it
from the raw response (for example resp.model_dump()) if the attribute is
missing. See
OpenAI compatibility.
What reasoning costs¶
Reasoning is not free, and it is not separately priced. It is generated tokens, reported separately and counted in the same total:
"usage": {
"prompt_tokens": 14,
"completion_tokens": 247,
"reasoning_tokens": 237,
"total_tokens": 261
}
In that response (a one-line greeting), 237 of the 247 generated tokens were
reasoning. reasoning_tokens is a breakdown of completion_tokens, not an
addition to it, so total_tokens is still prompt + completion. Budget
accordingly: a reasoning model can cost an order of magnitude more than a
standard one for the same visible output, which is the real reason to reserve it
for problems that need it.
Supported models¶
Reasoning is a property of the model. List the models your key can call with
GET /v1/models (see
Models and pricing) and check the model
card, or just make a call and look for reasoning_content in the response,
which is the definitive test.