Custom providers and local runtimes
Providers configured through models.providers: custom endpoints, base URLs, and local inference servers.
Providers via models.providers (custom/base URL)
Use models.providers (or models.json) to add custom providers or OpenAI/Anthropic-compatible proxies.
Many of the bundled provider plugins below already publish a default catalog. Use explicit models.providers.<id> entries only when you want to override the default base URL, headers, or model list.
Bundled and catalog-known routes take their compat capabilities from the owning provider plugin. A config compat block is for a custom provider/model or a different api/baseUrl route whose endpoint contract you have verified; see the custom-provider capability guide. Doctor removes legacy values that merely repeat the catalog and leaves divergent values visible for operator review.
A custom endpoint does not inherit the original catalog route's preferred Code Mode tier. With automatic Code Mode and no custom compat.codeMode declaration, the embedded runtime keeps its normal tool surface. Set compat.codeMode: "preferred" only after verifying Code Mode on that endpoint.
Provider-normalized aliases, such as an Anthropic endpoint's /v1 suffix, keep their canonical catalog capabilities.
Gateway model capability checks also read explicit models.providers.<id>.models[] metadata. If a custom or proxy model accepts images, set input: ["text", "image"] on that model so WebChat and node-origin attachment paths pass images as native model inputs instead of text-only media refs.
agents.defaults.models["provider/model"] controls aliases and per-model metadata for agents. It neither restricts overrides nor registers a new runtime model by itself. For custom provider models, also add models.providers.<provider>.models[] with at least the matching id; use agents.defaults.modelPolicy.allow separately when you want an override restriction.
Moonshot AI (Kimi)
Install @openclaw/moonshot-provider before onboarding. Add an explicit models.providers.moonshot entry only when you need to override the base URL or model metadata:
- Provider:
moonshot - Auth:
MOONSHOT_API_KEY - Example model:
moonshot/kimi-k3 - CLI:
openclaw onboard --auth-choice moonshot-api-keyoropenclaw onboard --auth-choice moonshot-api-key-cn
Kimi model IDs:
moonshot/kimi-k2.6moonshot/kimi-k3moonshot/kimi-k2.7-codemoonshot/kimi-k2.7-code-highspeedmoonshot/kimi-k2.5
{
agents: {
defaults: { model: { primary: "moonshot/kimi-k3" } },
},
models: {
mode: "merge",
providers: {
moonshot: {
baseUrl: "https://api.moonshot.ai/v1",
apiKey: "${MOONSHOT_API_KEY}",
api: "openai-completions",
models: [{ id: "kimi-k3", name: "Kimi K3" }],
},
},
},
}See Moonshot AI (Kimi + Kimi Coding) for the full setup guide.
Kimi Coding
Kimi Coding uses Moonshot AI's Anthropic-compatible endpoint:
- Provider:
kimi - Auth:
KIMI_API_KEY - Kimi K3:
kimi/k3(up to 1M, tier-gated) orkimi/k3-256k(256K, lower quota use) - Kimi Code:
kimi/kimi-for-coding - Kimi Code HighSpeed:
kimi/kimi-for-coding-highspeed
{
env: { vars: { KIMI_API_KEY: "sk-..." } },
agents: {
defaults: { model: { primary: "kimi/kimi-for-coding" } },
},
}Kimi K3 uses adaptive thinking. --thinking minimal|low selects low effort,
--thinking medium|high|adaptive selects high effort, and --thinking xhigh|max
selects max effort. Catalog pricing is $3/MTok input, $15/MTok output, and
$0.30/MTok cache reads. Legacy kimi/kimi-code and kimi/k2p5 remain
accepted as compatibility model ids and normalize to Kimi's stable API model
id; the previously published kimi/k3[1m] ref normalizes to kimi/k3 for
existing configs.
Volcano Engine (Doubao)
Volcano Engine (火山引擎) provides access to Doubao and other models in China.
- Provider:
volcengine(coding:volcengine-plan) - Auth:
VOLCANO_ENGINE_API_KEY - Example model:
volcengine-plan/ark-code-latest - CLI:
openclaw onboard --auth-choice volcengine-api-key
{
agents: {
defaults: { model: { primary: "volcengine-plan/ark-code-latest" } },
},
}Onboarding defaults to the coding surface, but the general volcengine/* catalog is registered at the same time.
In onboarding/configure model pickers, the Volcengine auth choice prefers both volcengine/* and volcengine-plan/* rows. If those models are not loaded yet, OpenClaw falls back to the unfiltered catalog instead of showing an empty provider-scoped picker.
Standard models
volcengine/doubao-seed-evolving(Doubao Seed Evolving)volcengine/doubao-seed-2-1-pro-260628(Doubao Seed 2.1 Pro)volcengine/doubao-seed-2-1-turbo-260628(Doubao Seed 2.1 Turbo)volcengine/glm-5-2-260617(GLM 5.2)volcengine/deepseek-v4-pro-260425(DeepSeek V4 Pro)volcengine/deepseek-v4-flash-260425(DeepSeek V4 Flash)
Coding models (volcengine-plan)
volcengine-plan/ark-code-latest(Ark Coding Plan)volcengine-plan/doubao-seed-2.1-turbo(Doubao Seed 2.1 Turbo)volcengine-plan/glm-5.2(GLM 5.2)volcengine-plan/deepseek-v4-pro(DeepSeek V4 Pro)volcengine-plan/deepseek-v4-flash(DeepSeek V4 Flash)
BytePlus (International)
BytePlus ARK provides access to the same models as Volcano Engine for international users.
- Plugin:
@openclaw/byteplus-provider - Provider:
byteplus(coding:byteplus-plan) - Auth:
BYTEPLUS_API_KEY - Example model:
byteplus-plan/ark-code-latest - CLI:
openclaw onboard --auth-choice byteplus-api-key
Install the official plugin and restart the Gateway:
openclaw plugins install @openclaw/byteplus-provider
openclaw gateway restart{
agents: {
defaults: { model: { primary: "byteplus-plan/ark-code-latest" } },
},
}Onboarding defaults to the coding surface, but the general byteplus/* catalog is registered at the same time.
In onboarding/configure model pickers, the BytePlus auth choice prefers both byteplus/* and byteplus-plan/* rows. If those models are not loaded yet, OpenClaw falls back to the unfiltered catalog instead of showing an empty provider-scoped picker.
Standard models
byteplus/dola-seed-2-1-turbo-260628(Dola Seed 2.1 Turbo)byteplus/seed-2-0-code-preview-260328(Seed 2.0 Code Preview)byteplus/glm-5-2-260617(GLM 5.2)byteplus/deepseek-v4-pro-260425(DeepSeek V4 Pro)byteplus/deepseek-v4-flash-260425(DeepSeek V4 Flash)
Coding models (byteplus-plan)
byteplus-plan/ark-code-latest(Ark Coding Plan)byteplus-plan/kimi-k2.5(Kimi K2.5 Coding)
Synthetic
Synthetic provides Anthropic-compatible models behind the synthetic provider:
- Provider:
synthetic - Auth:
SYNTHETIC_API_KEY - Example model:
synthetic/hf:MiniMaxAI/MiniMax-M3 - CLI:
openclaw onboard --auth-choice synthetic-api-key
{
agents: {
defaults: { model: { primary: "synthetic/hf:MiniMaxAI/MiniMax-M3" } },
},
models: {
mode: "merge",
providers: {
synthetic: {
baseUrl: "https://api.synthetic.new/anthropic",
apiKey: "${SYNTHETIC_API_KEY}",
api: "anthropic-messages",
models: [{ id: "hf:MiniMaxAI/MiniMax-M3", name: "MiniMax M3" }],
},
},
},
}MiniMax
MiniMax is configured via models.providers because it uses custom endpoints:
- MiniMax OAuth (Global):
--auth-choice minimax-global-oauth - MiniMax OAuth (CN):
--auth-choice minimax-cn-oauth - MiniMax API key (Global):
--auth-choice minimax-global-api - MiniMax API key (CN):
--auth-choice minimax-cn-api - Auth:
MINIMAX_API_KEYforminimax;MINIMAX_OAUTH_TOKENorMINIMAX_API_KEYforminimax-portal
See /providers/minimax for setup details, model options, and config snippets.
On MiniMax's Anthropic-compatible streaming path, OpenClaw disables thinking by default for the M2.x family unless you explicitly set it; MiniMax-M3 (and M3.x) stays on the provider's omitted/adaptive thinking path by default. /fast on rewrites MiniMax-M2.7 to MiniMax-M2.7-highspeed.
Plugin-owned capability split:
- Text/chat defaults stay on
minimax/MiniMax-M3 - Image generation is
minimax/image-01orminimax-portal/image-01 - Image understanding is plugin-owned
MiniMax-VL-01on both MiniMax auth paths - Web search stays on provider id
minimax
llama.cpp
The bundled llama-cpp plugin provides one local text provider with two setup choices:
- Managed local server installs and supervises a verified llama-server and local GGUF files.
- Existing llama-server connects to a server that you operate and discovers its models.
Install the plugin once for either path:
openclaw plugins install @openclaw/llama-cpp-providerBoth use llama-cpp/<model> references. See llama.cpp for setup,
discovery, authentication, and managed local embeddings.
llmman
llmman is configured via models.providers as an OpenAI-compatible local server. It pulls models as OCI artifacts and serves them through upstream llama-server, vllm, or mlx-lm, and can pair a local model with a hosted one under a single model id:
- Provider:
llmman(custom;api: "openai-completions") - Auth: none enforced; set
LLMMAN_API_KEY=llmman-localand useapiKey: "${LLMMAN_API_KEY}" - Default base URL:
http://127.0.0.1:17434/v1 - Example model:
llmman/qwen3.8 - Hybrid example:
llmman/llmman.hybrid/qwen3.8,openai/gpt-5.6-luna
llmman pull qwen3.8
llmman serve{
agents: {
defaults: { model: { primary: "llmman/qwen3.8" } },
},
models: {
providers: {
llmman: {
baseUrl: "http://127.0.0.1:17434/v1",
apiKey: "${LLMMAN_API_KEY}",
api: "openai-completions",
models: [
{ id: "qwen3.8", name: "Qwen3.8 (llmman)", reasoning: true, input: ["text", "image"] },
],
},
},
},
}See /providers/llmman for setup, hybrid local + hosted routing, vision, and troubleshooting.
LM Studio
LM Studio ships as a bundled provider plugin which uses the native API:
- Provider:
lmstudio - Auth:
LM_API_TOKEN - Default inference base URL:
http://localhost:1234/v1
Then set a model (replace with one of the IDs returned by http://localhost:1234/api/v1/models):
{
agents: {
defaults: { model: { primary: "lmstudio/openai/gpt-oss-20b" } },
},
}OpenClaw uses LM Studio's native /api/v1/models and /api/v1/models/load for discovery + auto-load, with /v1/chat/completions for inference by default. If you want LM Studio JIT loading, TTL, and auto-evict to own model lifecycle, set models.providers.lmstudio.params.preload: false. See /providers/lmstudio for setup and troubleshooting.
Ollama
Ollama ships as a bundled provider plugin and uses Ollama's native API:
- Provider:
ollama - Auth: None required (local server)
- Example model:
ollama/llama3.3 - Installation: https://ollama.com/download
# Install Ollama, then pull a model:
ollama pull llama3.3{
agents: {
defaults: { model: { primary: "ollama/llama3.3" } },
},
}Ollama is detected locally at http://127.0.0.1:11434 when you opt in with OLLAMA_API_KEY, and the bundled provider plugin adds Ollama directly to openclaw onboard and the model picker. See /providers/ollama for onboarding, cloud/local mode, and custom configuration.
vLLM
vLLM ships as a bundled provider plugin for local/self-hosted OpenAI-compatible servers:
- Provider:
vllm - Auth: Optional (depends on your server)
- Default base URL:
http://127.0.0.1:8000/v1
To opt in to auto-discovery locally (any value works if your server doesn't enforce auth):
export VLLM_API_KEY="vllm-local"Then set a model (replace with one of the IDs returned by /v1/models):
{
agents: {
defaults: { model: { primary: "vllm/your-model-id" } },
},
}See /providers/vllm for details.
SGLang
SGLang ships as a bundled provider plugin for fast self-hosted OpenAI-compatible servers:
- Provider:
sglang - Auth: Optional (depends on your server)
- Default base URL:
http://127.0.0.1:30000/v1
To opt in to auto-discovery locally (any value works if your server does not enforce auth):
export SGLANG_API_KEY="sglang-local"Then set a model (replace with one of the IDs returned by /v1/models):
{
agents: {
defaults: { model: { primary: "sglang/your-model-id" } },
},
}See /providers/sglang for details.
Local proxies (LM Studio, vLLM, LiteLLM, etc.)
Example (OpenAI-compatible):
{
agents: {
defaults: {
model: { primary: "lmstudio/my-local-model" },
models: { "lmstudio/my-local-model": { alias: "Local" } },
},
},
models: {
providers: {
lmstudio: {
baseUrl: "http://localhost:1234/v1",
apiKey: "${LM_API_TOKEN}",
api: "openai-completions",
timeoutSeconds: 300,
models: [
{
id: "my-local-model",
name: "Local Model",
reasoning: false,
input: ["text"],
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
contextWindow: 200000,
maxTokens: 8192,
},
],
},
},
},
}Default optional fields
For custom providers, reasoning, input, cost, contextWindow, and maxTokens are optional. When omitted, OpenClaw defaults to:
reasoning: falseinput: ["text"]cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }maxTokens: no fixed default. For OpenAI-compatible Completions, an unknown output limit omits bothmax_tokensandmax_completion_tokens, letting the provider apply its default.
Anthropic and Mistral requests preserve an explicit request output limit when the model's maxTokens is unknown. Anthropic manual thinking must fit within that request limit when no model output capacity is available.
An omitted contextWindow remains unset so authored native-window metadata is unambiguous. When neither discovery nor per-model context metadata is available, context-budget callers use the standard 200000-token fallback.
Recommended: set explicit values that match your proxy/model limits.
Model-selection metadata keeps capabilities tied to the API and endpoint that supplied them. A configured route change discards metadata from the previous route, while an explicit thinkingLevelMap is applied to the configured model.
Proxy-route shaping rules
- For
api: "openai-completions"on non-native endpoints (any non-emptybaseUrlwhose host is notapi.openai.com), OpenClaw forcescompat.supportsDeveloperRole: falseto avoid provider 400 errors for unsupporteddeveloperroles. - Runtime notices stay in conversation order as
developermessages when the route supports that role, or labeledusermessages otherwise. They never add a latersystemmessage, so strict chat templates can continue after a notice without changing the leading prompt or stored history. - Proxy-style OpenAI-compatible routes skip native OpenAI-only request shaping: no
service_tier, no Responsesstore, no Completionsstore, no prompt-cache hints, and no hidden OpenClaw attribution headers. - Custom
openai-completionsmodels markedreasoning: truesendreasoning_effortby default when thinking is enabled. Setcompat.supportsReasoningEffort: falseon the model if its endpoint rejects that field./think offomits it by default, leaving the server's own reasoning default in effect; see custom endpoint thinking for explicit off mappings and supported effort levels. - For OpenAI-compatible Completions proxies that need vendor-specific fields, set
agents.defaults.models["provider/model"].params.extra_body(orextraBody) to merge extra JSON into the outbound request body. - For vLLM chat-template controls, set
agents.defaults.models["provider/model"].params.chat_template_kwargs. The bundled vLLM plugin automatically sendsenable_thinking: falseandforce_nonempty_content: trueforvllm/nemotron-3-*when the session thinking level is off. - For slow local models or remote LAN/tailnet hosts, set
models.providers.<id>.timeoutSeconds. This extends provider model HTTP request handling, including connect, headers, body streaming, and the total guarded-fetch abort, without increasing the whole agent runtime timeout. The initial response-header wait uses the first-event budget: 120 seconds for cloud endpoints and 300 seconds for local/self-hosted endpoints, unless overridden by the provider timeout or a lower run ceiling. This does not add idle-gap policing to local streams after the provider accepts the request. Without an explicit provider timeout, TCP/TLS connection setup keeps its 10-second default independently of the longer streaming timeout, including when reconnecting a pooled transport. Ifagents.defaults.timeoutSecondsor a run-specific timeout is lower, raise that ceiling too; provider timeouts cannot extend the whole run. - Model provider HTTP calls allow Surge, Clash, and sing-box fake-IP DNS answers in
198.18.0.0/15andfc00::/7only for the configured providerbaseUrlhostname. Custom/local provider endpoints also trust that exact configuredscheme://host:portorigin for guarded model requests, including loopback, LAN, and tailnet hosts. This is not a new config option; thebaseUrlyou configure extends the request policy only for that origin. Fake-IP hostname allowance and exact-origin trust are independent mechanisms. Other private, loopback, link-local, metadata, local-use NAT64 (64:ff9b:1::/48) destinations, and different ports still require an explicitmodels.providers.<id>.request.allowPrivateNetwork: trueopt-in. Setmodels.providers.<id>.request.allowPrivateNetwork: falseto opt out of the exact-origin trust. - If
baseUrlis empty/omitted, OpenClaw keeps the default OpenAI behavior (which resolves toapi.openai.com). - For safety, an explicit
compat.supportsDeveloperRole: trueis still overridden on non-nativeopenai-completionsendpoints. - For
api: "anthropic-messages"on non-direct endpoints (any provider other than canonicalanthropic, or a custommodels.providers.anthropic.baseUrlwhose host is not a publicapi.anthropic.comendpoint), OpenClaw suppresses implicit Anthropic beta headers such asclaude-code-20250219,interleaved-thinking-2025-05-14, and OAuth markers, so custom Anthropic-compatible proxies do not reject unsupported beta flags. Setmodels.providers.<id>.headers["anthropic-beta"]explicitly if your proxy needs specific beta features.