Structured output¶
Force a model's response to conform to a schema (for example, valid JSON that
matches a JSON Schema you provide) so you can parse it reliably. Mango Inference
uses the OpenAI-compatible response_format parameter.
When to use it¶
Use structured output when a downstream system consumes the response and you need guaranteed shape (JSON), not free-form prose.
JSON mode¶
resp = client.chat.completions.create(
model="zai-org/GLM-5.3",
messages=[{"role": "user", "content": "List 3 fruits as JSON."}],
response_format={"type": "json_object"},
)
JSON Schema¶
Constrain the output to a specific schema:
resp = client.chat.completions.create(
model="zai-org/GLM-5.3",
messages=[{"role": "user", "content": "Extract the person's name and age."}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "person",
"schema": {
"type": "object",
"properties": {"name": {"type": "string"}, "age": {"type": "integer"}},
"required": ["name", "age"],
},
},
},
)
The response arrives the same way as any other completion. The content is a JSON string, so parse it yourself:
import json
data = json.loads(resp.choices[0].message.content)
print(data["name"], data["age"])
Which modes work¶
response_format is passed through to the model server rather than validated
by Mango Inference, so the modes available to you are the ones your chosen model
and its serving stack implement. json_object and json_schema are the two
these docs cover; other modes some servers expose (grammar or regex
constraints) are not documented here and should not be assumed.
Verify with a real call before depending on a mode. A mode that is not enforced
fails quietly and expensively: you get plausible free-form text back, and the
first thing that notices is json.loads in production.
Ask for the shape in the prompt too
Even with json_schema set, describing the expected fields in the prompt
produces better-populated output. The schema constrains the shape; it does
not tell the model what you meant by name.
Supported models¶
Structured output depends on the model, not the platform. See Models and pricing for how to list the models your key can call, and check the model card before relying on it.
If you need a guarantee rather than a strong tendency, validate the parsed result against your own schema and retry on failure. That is worth doing even where the mode is enforced.
Related¶
- Tool calling
- Chat Completions reference: the
response_formatparameter.