# Custom providers and local runtimes

> Providers configured through models.providers: custom endpoints, base URLs, and local inference servers.

- 网址：https://funcoding.ai/agents/openclaw/concepts/model-providers/custom-providers/
- 来源：OpenClaw 官方文档原文（英文），MIT 许可，同步于 2026-10-11
- 官方原文：https://docs.openclaw.ai/zh-CN/concepts/model-providers/custom-providers

---
## Providers via `models.providers` (custom/base URL)

Use `models.providers` (or `models.json`) to add **custom** providers or OpenAI/Anthropic-compatible proxies.

Many of the bundled provider plugins below already publish a default catalog. Use explicit `models.providers.<id>` entries only when you want to override the default base URL, headers, or model list.

Bundled and catalog-known routes take their `compat` capabilities from the owning provider plugin. A config `compat` block is for a custom provider/model or a different `api`/`baseUrl` route whose endpoint contract you have verified; see the [custom-provider capability guide](https://funcoding.ai/agents/openclaw/gateway/config-tools/#custom-provider-capability-declarations). Doctor removes legacy values that merely repeat the catalog and leaves divergent values visible for operator review.

A custom endpoint does not inherit the original catalog route's preferred Code Mode tier. With automatic Code Mode and no custom `compat.codeMode` declaration, the embedded runtime keeps its normal tool surface. Set `compat.codeMode: "preferred"` only after verifying Code Mode on that endpoint.

Provider-normalized aliases, such as an Anthropic endpoint's `/v1` suffix, keep their canonical catalog capabilities.

Gateway model capability checks also read explicit `models.providers.<id>.models[]` metadata. If a custom or proxy model accepts images, set `input: ["text", "image"]` on that model so WebChat and node-origin attachment paths pass images as native model inputs instead of text-only media refs.

`agents.defaults.models["provider/model"]` controls aliases and per-model metadata for agents. It neither restricts overrides nor registers a new runtime model by itself. For custom provider models, also add `models.providers.<provider>.models[]` with at least the matching `id`; use `agents.defaults.modelPolicy.allow` separately when you want an override restriction.

### Moonshot AI (Kimi)

Install `@openclaw/moonshot-provider` before onboarding. Add an explicit `models.providers.moonshot` entry only when you need to override the base URL or model metadata:

- Provider: `moonshot`
- Auth: `MOONSHOT_API_KEY`
- Example model: `moonshot/kimi-k3`
- CLI: `openclaw onboard --auth-choice moonshot-api-key` or `openclaw onboard --auth-choice moonshot-api-key-cn`

Kimi model IDs:

[//]: # "moonshot-kimi-k2-model-refs:start"

- `moonshot/kimi-k2.6`
- `moonshot/kimi-k3`
- `moonshot/kimi-k2.7-code`
- `moonshot/kimi-k2.7-code-highspeed`
- `moonshot/kimi-k2.5`

[//]: # "moonshot-kimi-k2-model-refs:end"

```json5
{
  agents: {
    defaults: { model: { primary: "moonshot/kimi-k3" } },
  },
  models: {
    mode: "merge",
    providers: {
      moonshot: {
        baseUrl: "https://api.moonshot.ai/v1",
        apiKey: "${MOONSHOT_API_KEY}",
        api: "openai-completions",
        models: [{ id: "kimi-k3", name: "Kimi K3" }],
      },
    },
  },
}
```

See [Moonshot AI (Kimi + Kimi Coding)](https://funcoding.ai/agents/openclaw/providers/moonshot/) for the full setup guide.

### Kimi Coding

Kimi Coding uses Moonshot AI's Anthropic-compatible endpoint:

- Provider: `kimi`
- Auth: `KIMI_API_KEY`
- Kimi K3: `kimi/k3` (up to 1M, tier-gated) or `kimi/k3-256k` (256K, lower quota use)
- Kimi Code: `kimi/kimi-for-coding`
- Kimi Code HighSpeed: `kimi/kimi-for-coding-highspeed`

```json5
{
  env: { vars: { KIMI_API_KEY: "sk-..." } },
  agents: {
    defaults: { model: { primary: "kimi/kimi-for-coding" } },
  },
}
```

Kimi K3 uses adaptive thinking. `--thinking minimal|low` selects low effort,
`--thinking medium|high|adaptive` selects high effort, and `--thinking xhigh|max`
selects max effort. Catalog pricing is $3/MTok input, $15/MTok output, and
$0.30/MTok cache reads. Legacy `kimi/kimi-code` and `kimi/k2p5` remain
accepted as compatibility model ids and normalize to Kimi's stable API model
id; the previously published `kimi/k3[1m]` ref normalizes to `kimi/k3` for
existing configs.

### Volcano Engine (Doubao)

Volcano Engine (火山引擎) provides access to Doubao and other models in China.

- Provider: `volcengine` (coding: `volcengine-plan`)
- Auth: `VOLCANO_ENGINE_API_KEY`
- Example model: `volcengine-plan/ark-code-latest`
- CLI: `openclaw onboard --auth-choice volcengine-api-key`

```json5
{
  agents: {
    defaults: { model: { primary: "volcengine-plan/ark-code-latest" } },
  },
}
```

Onboarding defaults to the coding surface, but the general `volcengine/*` catalog is registered at the same time.

In onboarding/configure model pickers, the Volcengine auth choice prefers both `volcengine/*` and `volcengine-plan/*` rows. If those models are not loaded yet, OpenClaw falls back to the unfiltered catalog instead of showing an empty provider-scoped picker.

**Standard models**

- `volcengine/doubao-seed-evolving` (Doubao Seed Evolving)
- `volcengine/doubao-seed-2-1-pro-260628` (Doubao Seed 2.1 Pro)
- `volcengine/doubao-seed-2-1-turbo-260628` (Doubao Seed 2.1 Turbo)
- `volcengine/glm-5-2-260617` (GLM 5.2)
- `volcengine/deepseek-v4-pro-260425` (DeepSeek V4 Pro)
- `volcengine/deepseek-v4-flash-260425` (DeepSeek V4 Flash)

**Coding models (volcengine-plan)**

- `volcengine-plan/ark-code-latest` (Ark Coding Plan)
- `volcengine-plan/doubao-seed-2.1-turbo` (Doubao Seed 2.1 Turbo)
- `volcengine-plan/glm-5.2` (GLM 5.2)
- `volcengine-plan/deepseek-v4-pro` (DeepSeek V4 Pro)
- `volcengine-plan/deepseek-v4-flash` (DeepSeek V4 Flash)

### BytePlus (International)

BytePlus ARK provides access to the same models as Volcano Engine for international users.

- Plugin: `@openclaw/byteplus-provider`
- Provider: `byteplus` (coding: `byteplus-plan`)
- Auth: `BYTEPLUS_API_KEY`
- Example model: `byteplus-plan/ark-code-latest`
- CLI: `openclaw onboard --auth-choice byteplus-api-key`

Install the official plugin and restart the Gateway:

```bash
openclaw plugins install @openclaw/byteplus-provider
openclaw gateway restart
```

```json5
{
  agents: {
    defaults: { model: { primary: "byteplus-plan/ark-code-latest" } },
  },
}
```

Onboarding defaults to the coding surface, but the general `byteplus/*` catalog is registered at the same time.

In onboarding/configure model pickers, the BytePlus auth choice prefers both `byteplus/*` and `byteplus-plan/*` rows. If those models are not loaded yet, OpenClaw falls back to the unfiltered catalog instead of showing an empty provider-scoped picker.

**Standard models**

- `byteplus/dola-seed-2-1-turbo-260628` (Dola Seed 2.1 Turbo)
- `byteplus/seed-2-0-code-preview-260328` (Seed 2.0 Code Preview)
- `byteplus/glm-5-2-260617` (GLM 5.2)
- `byteplus/deepseek-v4-pro-260425` (DeepSeek V4 Pro)
- `byteplus/deepseek-v4-flash-260425` (DeepSeek V4 Flash)

**Coding models (byteplus-plan)**

- `byteplus-plan/ark-code-latest` (Ark Coding Plan)
- `byteplus-plan/kimi-k2.5` (Kimi K2.5 Coding)

### Synthetic

Synthetic provides Anthropic-compatible models behind the `synthetic` provider:

- Provider: `synthetic`
- Auth: `SYNTHETIC_API_KEY`
- Example model: `synthetic/hf:MiniMaxAI/MiniMax-M3`
- CLI: `openclaw onboard --auth-choice synthetic-api-key`

```json5
{
  agents: {
    defaults: { model: { primary: "synthetic/hf:MiniMaxAI/MiniMax-M3" } },
  },
  models: {
    mode: "merge",
    providers: {
      synthetic: {
        baseUrl: "https://api.synthetic.new/anthropic",
        apiKey: "${SYNTHETIC_API_KEY}",
        api: "anthropic-messages",
        models: [{ id: "hf:MiniMaxAI/MiniMax-M3", name: "MiniMax M3" }],
      },
    },
  },
}
```

### MiniMax

MiniMax is configured via `models.providers` because it uses custom endpoints:

- MiniMax OAuth (Global): `--auth-choice minimax-global-oauth`
- MiniMax OAuth (CN): `--auth-choice minimax-cn-oauth`
- MiniMax API key (Global): `--auth-choice minimax-global-api`
- MiniMax API key (CN): `--auth-choice minimax-cn-api`
- Auth: `MINIMAX_API_KEY` for `minimax`; `MINIMAX_OAUTH_TOKEN` or `MINIMAX_API_KEY` for `minimax-portal`

See [/providers/minimax](https://funcoding.ai/agents/openclaw/providers/minimax/) for setup details, model options, and config snippets.

<div class="callout callout-note">

On MiniMax's Anthropic-compatible streaming path, OpenClaw disables thinking by default for the M2.x family unless you explicitly set it; MiniMax-M3 (and M3.x) stays on the provider's omitted/adaptive thinking path by default. `/fast on` rewrites `MiniMax-M2.7` to `MiniMax-M2.7-highspeed`.

</div>

Plugin-owned capability split:

- Text/chat defaults stay on `minimax/MiniMax-M3`
- Image generation is `minimax/image-01` or `minimax-portal/image-01`
- Image understanding is plugin-owned `MiniMax-VL-01` on both MiniMax auth paths
- Web search stays on provider id `minimax`

### llama.cpp

The bundled `llama-cpp` plugin provides one local text provider with two setup choices:

- **Managed local server** installs and supervises a verified llama-server and local GGUF files.
- **Existing llama-server** connects to a server that you operate and discovers its models.

Install the plugin once for either path:

```bash
openclaw plugins install @openclaw/llama-cpp-provider
```

Both use `llama-cpp/<model>` references. See [llama.cpp](https://funcoding.ai/agents/openclaw/plugins/llama-cpp/) for setup,
discovery, authentication, and managed local embeddings.

### llmman

llmman is configured via `models.providers` as an OpenAI-compatible local server. It pulls models as OCI artifacts and serves them through upstream `llama-server`, `vllm`, or `mlx-lm`, and can pair a local model with a hosted one under a single model id:

- Provider: `llmman` (custom; `api: "openai-completions"`)
- Auth: none enforced; set `LLMMAN_API_KEY=llmman-local` and use `apiKey: "${LLMMAN_API_KEY}"`
- Default base URL: `http://127.0.0.1:17434/v1`
- Example model: `llmman/qwen3.8`
- Hybrid example: `llmman/llmman.hybrid/qwen3.8,openai/gpt-5.6-luna`

```bash
llmman pull qwen3.8
llmman serve
```

```json5
{
  agents: {
    defaults: { model: { primary: "llmman/qwen3.8" } },
  },
  models: {
    providers: {
      llmman: {
        baseUrl: "http://127.0.0.1:17434/v1",
        apiKey: "${LLMMAN_API_KEY}",
        api: "openai-completions",
        models: [
          { id: "qwen3.8", name: "Qwen3.8 (llmman)", reasoning: true, input: ["text", "image"] },
        ],
      },
    },
  },
}
```

See [/providers/llmman](https://funcoding.ai/agents/openclaw/providers/llmman/) for setup, hybrid local + hosted routing, vision, and troubleshooting.

### LM Studio

LM Studio ships as a bundled provider plugin which uses the native API:

- Provider: `lmstudio`
- Auth: `LM_API_TOKEN`
- Default inference base URL: `http://localhost:1234/v1`

Then set a model (replace with one of the IDs returned by `http://localhost:1234/api/v1/models`):

```json5
{
  agents: {
    defaults: { model: { primary: "lmstudio/openai/gpt-oss-20b" } },
  },
}
```

OpenClaw uses LM Studio's native `/api/v1/models` and `/api/v1/models/load` for discovery + auto-load, with `/v1/chat/completions` for inference by default. If you want LM Studio JIT loading, TTL, and auto-evict to own model lifecycle, set `models.providers.lmstudio.params.preload: false`. See [/providers/lmstudio](https://funcoding.ai/agents/openclaw/providers/lmstudio/) for setup and troubleshooting.

### Ollama

Ollama ships as a bundled provider plugin and uses Ollama's native API:

- Provider: `ollama`
- Auth: None required (local server)
- Example model: `ollama/llama3.3`
- Installation: [https://ollama.com/download](https://ollama.com/download)

```bash
# Install Ollama, then pull a model:
ollama pull llama3.3
```

```json5
{
  agents: {
    defaults: { model: { primary: "ollama/llama3.3" } },
  },
}
```

Ollama is detected locally at `http://127.0.0.1:11434` when you opt in with `OLLAMA_API_KEY`, and the bundled provider plugin adds Ollama directly to `openclaw onboard` and the model picker. See [/providers/ollama](https://funcoding.ai/agents/openclaw/providers/ollama/) for onboarding, cloud/local mode, and custom configuration.

### vLLM

vLLM ships as a bundled provider plugin for local/self-hosted OpenAI-compatible servers:

- Provider: `vllm`
- Auth: Optional (depends on your server)
- Default base URL: `http://127.0.0.1:8000/v1`

To opt in to auto-discovery locally (any value works if your server doesn't enforce auth):

```bash
export VLLM_API_KEY="vllm-local"
```

Then set a model (replace with one of the IDs returned by `/v1/models`):

```json5
{
  agents: {
    defaults: { model: { primary: "vllm/your-model-id" } },
  },
}
```

See [/providers/vllm](https://funcoding.ai/agents/openclaw/providers/vllm/) for details.

### SGLang

SGLang ships as a bundled provider plugin for fast self-hosted OpenAI-compatible servers:

- Provider: `sglang`
- Auth: Optional (depends on your server)
- Default base URL: `http://127.0.0.1:30000/v1`

To opt in to auto-discovery locally (any value works if your server does not enforce auth):

```bash
export SGLANG_API_KEY="sglang-local"
```

Then set a model (replace with one of the IDs returned by `/v1/models`):

```json5
{
  agents: {
    defaults: { model: { primary: "sglang/your-model-id" } },
  },
}
```

See [/providers/sglang](https://funcoding.ai/agents/openclaw/providers/sglang/) for details.

### Local proxies (LM Studio, vLLM, LiteLLM, etc.)

Example (OpenAI-compatible):

```json5
{
  agents: {
    defaults: {
      model: { primary: "lmstudio/my-local-model" },
      models: { "lmstudio/my-local-model": { alias: "Local" } },
    },
  },
  models: {
    providers: {
      lmstudio: {
        baseUrl: "http://localhost:1234/v1",
        apiKey: "${LM_API_TOKEN}",
        api: "openai-completions",
        timeoutSeconds: 300,
        models: [
          {
            id: "my-local-model",
            name: "Local Model",
            reasoning: false,
            input: ["text"],
            cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
            contextWindow: 200000,
            maxTokens: 8192,
          },
        ],
      },
    },
  },
}
```

<details>
<summary>Default optional fields</summary>

For custom providers, `reasoning`, `input`, `cost`, `contextWindow`, and `maxTokens` are optional. When omitted, OpenClaw defaults to:

- `reasoning: false`
- `input: ["text"]`
- `cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }`
- `maxTokens`: no fixed default. For OpenAI-compatible Completions, an unknown output limit omits both `max_tokens` and `max_completion_tokens`, letting the provider apply its default.

Anthropic and Mistral requests preserve an explicit request output limit when the model's `maxTokens` is unknown. Anthropic manual thinking must fit within that request limit when no model output capacity is available.

An omitted `contextWindow` remains unset so authored native-window metadata is unambiguous. When neither discovery nor per-model context metadata is available, context-budget callers use the standard `200000`-token fallback.

Recommended: set explicit values that match your proxy/model limits.

Model-selection metadata keeps capabilities tied to the API and endpoint that supplied them. A configured route change discards metadata from the previous route, while an explicit `thinkingLevelMap` is applied to the configured model.

</details>

<details>
<summary>Proxy-route shaping rules</summary>

- For `api: "openai-completions"` on non-native endpoints (any non-empty `baseUrl` whose host is not `api.openai.com`), OpenClaw forces `compat.supportsDeveloperRole: false` to avoid provider 400 errors for unsupported `developer` roles.
- Runtime notices stay in conversation order as `developer` messages when the route supports that role, or labeled `user` messages otherwise. They never add a later `system` message, so strict chat templates can continue after a notice without changing the leading prompt or stored history.
- Proxy-style OpenAI-compatible routes skip native OpenAI-only request shaping: no `service_tier`, no Responses `store`, no Completions `store`, no prompt-cache hints, and no hidden OpenClaw attribution headers.
- Custom `openai-completions` models marked `reasoning: true` send `reasoning_effort` by default when thinking is enabled. Set `compat.supportsReasoningEffort: false` on the model if its endpoint rejects that field. `/think off` omits it by default, leaving the server's own reasoning default in effect; see [custom endpoint thinking](https://funcoding.ai/agents/openclaw/tools/thinking/#custom-openai-compatible-endpoints) for explicit off mappings and supported effort levels.
- For OpenAI-compatible Completions proxies that need vendor-specific fields, set `agents.defaults.models["provider/model"].params.extra_body` (or `extraBody`) to merge extra JSON into the outbound request body.
- For vLLM chat-template controls, set `agents.defaults.models["provider/model"].params.chat_template_kwargs`. The bundled vLLM plugin automatically sends `enable_thinking: false` and `force_nonempty_content: true` for `vllm/nemotron-3-*` when the session thinking level is off.
- For slow local models or remote LAN/tailnet hosts, set `models.providers.<id>.timeoutSeconds`. This extends provider model HTTP request handling, including connect, headers, body streaming, and the total guarded-fetch abort, without increasing the whole agent runtime timeout. The initial response-header wait uses the first-event budget: 120 seconds for cloud endpoints and 300 seconds for local/self-hosted endpoints, unless overridden by the provider timeout or a lower run ceiling. This does not add idle-gap policing to local streams after the provider accepts the request. Without an explicit provider timeout, TCP/TLS connection setup keeps its 10-second default independently of the longer streaming timeout, including when reconnecting a pooled transport. If `agents.defaults.timeoutSeconds` or a run-specific timeout is lower, raise that ceiling too; provider timeouts cannot extend the whole run.
- Model provider HTTP calls allow Surge, Clash, and sing-box fake-IP DNS answers in `198.18.0.0/15` and `fc00::/7` only for the configured provider `baseUrl` hostname. Custom/local provider endpoints also trust that exact configured `scheme://host:port` origin for guarded model requests, including loopback, LAN, and tailnet hosts. This is not a new config option; the `baseUrl` you configure extends the request policy only for that origin. Fake-IP hostname allowance and exact-origin trust are independent mechanisms. Other private, loopback, link-local, metadata, local-use NAT64 (`64:ff9b:1::/48`) destinations, and different ports still require an explicit `models.providers.<id>.request.allowPrivateNetwork: true` opt-in. Set `models.providers.<id>.request.allowPrivateNetwork: false` to opt out of the exact-origin trust.
- If `baseUrl` is empty/omitted, OpenClaw keeps the default OpenAI behavior (which resolves to `api.openai.com`).
- For safety, an explicit `compat.supportsDeveloperRole: true` is still overridden on non-native `openai-completions` endpoints.
- For `api: "anthropic-messages"` on non-direct endpoints (any provider other than canonical `anthropic`, or a custom `models.providers.anthropic.baseUrl` whose host is not a public `api.anthropic.com` endpoint), OpenClaw suppresses implicit Anthropic beta headers such as `claude-code-20250219`, `interleaved-thinking-2025-05-14`, and OAuth markers, so custom Anthropic-compatible proxies do not reject unsupported beta flags. Set `models.providers.<id>.headers["anthropic-beta"]` explicitly if your proxy needs specific beta features.

</details>
