Skip to content

Claude Code

Point Claude Code at Mango Inference's Anthropic-compatible endpoint by setting two environment variables: the base URL and your API key.

Prerequisites

  • A Mango Inference API key.
  • Claude Code installed.

1. Set environment variables

export ANTHROPIC_BASE_URL="https://api.mangoboost.io"
export ANTHROPIC_API_KEY="<your-api-key>"

Claude Code sends ANTHROPIC_API_KEY in the x-api-key header, which matches how Mango Inference authenticates /v1/messages requests. See Authentication. Set the base URL without a /v1 suffix; Claude Code appends /v1/messages itself.

2. Point at a Mango Inference model

export ANTHROPIC_DEFAULT_OPUS_MODEL="zai-org/GLM-5.3"
export ANTHROPIC_DEFAULT_SONNET_MODEL="zai-org/GLM-5.3"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="zai-org/GLM-5.3"

Whatever string you put in these variables is sent through as the model field, so any ID from your model list works. There is nothing to translate for Mango Inference. Set all three: Claude Code picks a tier per task, and a tier left pointing at an Anthropic model name will fail once it is selected.

Which variable names Claude Code reads is Claude Code's business, not Mango Inference's, and the set has changed between releases (the older ANTHROPIC_MODEL / ANTHROPIC_SMALL_FAST_MODEL pair predates the per-tier variables above). If a version ignores these, check Claude Code's own documentation for the current names. The Mango Inference side is unchanged either way.

3. Set the context length

A model name can carry a context length specifier in square brackets, and Mango Inference reads it as part of the model ID. zai-org/GLM-5.2-FP8 supports context lengths up to 1 million tokens, written [1m]:

export ANTHROPIC_DEFAULT_OPUS_MODEL="zai-org/GLM-5.2-FP8[1m]"
export ANTHROPIC_DEFAULT_SONNET_MODEL="zai-org/GLM-5.2-FP8[1m]"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="zai-org/GLM-5.2-FP8[1m]"

The suffix is part of the model string you pass in the ANTHROPIC_DEFAULT_*_MODEL variables (there is no separate setting for it), and it is how you control how much context the model will accept. Query GET /v1/models for the available model IDs and the context lengths each one supports; see also Models and pricing.

4. Run Claude Code

claude

Claude Code starts and routes requests through Mango Inference instead of Anthropic's API.

What to expect

Mango Inference implements the Messages endpoint, not the whole Anthropic platform, and Claude Code uses more of that platform than a plain chat client does. Two consequences worth knowing before you start:

  • /v1/messages/count_tokens is served. Mango Inference relays the token count from the upstream engine, so Claude Code's context estimates work. The call consumes the key's RPM slot but is not billed (no tokens are generated).
  • Prompt caching and extended thinking are Anthropic API features, not parts of the Messages request shape that a compatible gateway necessarily implements. Whether they take effect depends on the model and the gateway; neither is verified here. If they are not honoured, the practical effect is cost and latency, not failure.

Everything Claude Code does beyond that (tool use, multi-turn editing, streaming) rides on ordinary /v1/messages calls, and depends on the model you point it at being good at agentic tool use. A model that is weak at tool calling produces a Claude Code that talks about editing files instead of editing them; that is a model-choice problem, not a configuration one. See Tool calling.