Claude Code¶
Point Claude Code at Mango Inference's Anthropic-compatible endpoint by setting two environment variables: the base URL and your API key.
Prerequisites¶
- A Mango Inference API key.
- Claude Code installed.
1. Set environment variables¶
export ANTHROPIC_BASE_URL="https://api.mangoboost.io"
export ANTHROPIC_API_KEY="<your-api-key>"
Claude Code sends ANTHROPIC_API_KEY in the x-api-key header, which matches
how Mango Inference authenticates /v1/messages requests. See
Authentication. Set the base URL without
a /v1 suffix; Claude Code appends /v1/messages itself.
2. Point at a Mango Inference model¶
export ANTHROPIC_DEFAULT_OPUS_MODEL="zai-org/GLM-5.3"
export ANTHROPIC_DEFAULT_SONNET_MODEL="zai-org/GLM-5.3"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="zai-org/GLM-5.3"
Whatever string you put in these variables is sent through as the model
field, so any ID from your model list
works. There is nothing to translate for Mango Inference. Set all three:
Claude Code picks a tier per task, and a tier left pointing at an Anthropic
model name will fail once it is selected.
Which variable names Claude Code reads is Claude Code's business, not
Mango Inference's, and the set has changed between releases (the older
ANTHROPIC_MODEL / ANTHROPIC_SMALL_FAST_MODEL pair predates the per-tier
variables above). If a version ignores these, check
Claude Code's own documentation
for the current names. The Mango Inference side is unchanged either way.
3. Set the context length¶
A model name can carry a context length specifier in square brackets, and
Mango Inference reads it as part of the model ID. zai-org/GLM-5.2-FP8 supports
context lengths up to 1 million tokens, written [1m]:
export ANTHROPIC_DEFAULT_OPUS_MODEL="zai-org/GLM-5.2-FP8[1m]"
export ANTHROPIC_DEFAULT_SONNET_MODEL="zai-org/GLM-5.2-FP8[1m]"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="zai-org/GLM-5.2-FP8[1m]"
The suffix is part of the model string you pass in the
ANTHROPIC_DEFAULT_*_MODEL variables (there is no separate setting for it),
and it is how you control how much context the model will accept. Query
GET /v1/models for the available model IDs and the context lengths each one
supports; see also
Models and pricing.
4. Run Claude Code¶
claude
Claude Code starts and routes requests through Mango Inference instead of Anthropic's API.
What to expect¶
Mango Inference implements the Messages endpoint, not the whole Anthropic platform, and Claude Code uses more of that platform than a plain chat client does. Two consequences worth knowing before you start:
/v1/messages/count_tokensis served. Mango Inference relays the token count from the upstream engine, so Claude Code's context estimates work. The call consumes the key's RPM slot but is not billed (no tokens are generated).- Prompt caching and extended thinking are Anthropic API features, not parts of the Messages request shape that a compatible gateway necessarily implements. Whether they take effect depends on the model and the gateway; neither is verified here. If they are not honoured, the practical effect is cost and latency, not failure.
Everything Claude Code does beyond that (tool use, multi-turn editing,
streaming) rides on ordinary /v1/messages calls, and depends on the model
you point it at being good at agentic tool use. A model that is weak at tool
calling produces a Claude Code that talks about editing files instead of
editing them; that is a model-choice problem, not a configuration one. See
Tool calling.