跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

Ollama advanced configuration

Context windows, thinking control, costs, embeddings, and streaming for Ollama

Advanced configuration

Native Ollama requests keep exact session identities and runtime facts in the first user message, after the shared system and tool prefix. This allows local prompt caches to reuse that prefix across equivalent subagent spawns while preserving earlier message bytes on follow-up turns. Cache reuse still depends on the model template, available cache slots, and unchanged instructions/tools.

Legacy OpenAI-compatible mode

Tool calling is not reliable in this mode. Use it only when a proxy needs OpenAI format and you do not depend on native tool calling.

Set api: "openai-completions" explicitly for a proxy behind /v1/chat/completions:

{
  models: {
    providers: {
      ollama: {
        baseUrl: "http://ollama-host:11434/v1",
        api: "openai-completions",
        injectNumCtxForOpenAICompat: true, // default: true
        apiKey: "ollama-local",
        models: [...]
      }
    }
  }
}

This mode may not support streaming and tool calling simultaneously; you may need params: { streaming: false } on the model in models.providers.ollama.models. You can also set it under agents.defaults.models["ollama/<model>"].params. For openai-completions, this sends stream: false upstream and delivers the completed reply with its usage, reasoning, and tool calls. Streaming remains enabled when unset; this setting does not change native Ollama /api/chat requests.

OpenClaw injects options.num_ctx by default in this mode so Ollama does not silently fall back to a 4096-token context. If your proxy rejects unknown options fields, disable it:

{
  models: {
    providers: {
      ollama: {
        baseUrl: "http://ollama-host:11434/v1",
        api: "openai-completions",
        injectNumCtxForOpenAICompat: false,
        apiKey: "ollama-local",
        models: [...]
      }
    }
  }
}
Context windows

For auto-discovered models, OpenClaw uses the context window /api/show reports, including larger PARAMETER num_ctx values from custom Modelfiles; otherwise it falls back to OpenClaw's default Ollama context window.

Per-model contextWindow declares native window metadata, and per-model contextTokens caps active input. Provider-level maxTokens remains an output-token default; a model entry can override it. Native /api/chat requests set options.num_ctx from a positive params.num_ctx first, then from the effective model contextTokens when present. Local discovery normally caps contextTokens at 32,768 (or the model's smaller native window), so OpenClaw can override a smaller Modelfile context even without an explicit params.num_ctx. Invalid, zero, negative, or non-finite params.num_ctx values are ignored. Only when neither value is available does Ollama choose its own model, Modelfile, OLLAMA_CONTEXT_LENGTH, or VRAM-based default; the native adapter does not fall back directly to the advertised contextWindow. After upgrading an older configuration, run openclaw doctor --fix. Doctor preserves current contextTokens caps without creating a stronger model or provider num_ctx pin; uncapped legacy native entries still migrate their older context budgets. Existing explicit params.num_ctx values remain authoritative, including pins an older Doctor already wrote. Review or remove an oversized existing pin to let contextTokens drive the request again. Use params.num_ctx to override the native request context explicitly. The OpenAI-compatible adapter still injects options.num_ctx by default from params.num_ctx, then the matching model entry's contextTokens or contextWindow; disable with injectNumCtxForOpenAICompat: false if the upstream rejects options.

Native model entries also accept common Ollama runtime options under params, forwarded as native /api/chat options: num_keep, seed, num_predict, top_k, top_p, min_p, typical_p, repeat_last_n, temperature, repeat_penalty, presence_penalty, frequency_penalty, stop, num_batch, num_gpu, main_gpu, use_mmap, and num_thread. Runtime sampling controls (temperature, topP, frequencyPenalty, presencePenalty, and seed) override the matching model defaults, including explicit zero values. The Gateway's Chat Completions API maps top_p, frequency_penalty, and presence_penalty to these controls. With temperature: 0, OpenClaw still normalizes top_p to 1 for greedy sampling after applying overrides. A few keys (format, keep_alive, truncate, shift) are forwarded as top-level request fields instead of nested options. Local native chat requests default to truncate: false and shift: false, so supporting servers reject overflowing input instead of silently dropping history. OpenClaw then attempts compaction and retries, or reports the failure. Generation that fills the window can still produce a labeled partial reply. This behavior is verified with Ollama 0.33.3; older servers may ignore the fields. Explicit per-model values override these defaults. Hosted models and the OpenAI-compatible endpoint keep their existing behavior. OpenClaw only forwards these Ollama request keys, so runtime-only params such as streaming are never sent to Ollama. Use params.think (or params.thinking) to set top-level think; false disables API-level thinking for Qwen-style thinking models.

{
  models: {
    providers: {
      ollama: {
        models: [
          {
            id: "llama3.3",
            contextWindow: 131072,
            contextTokens: 32768,
            maxTokens: 65536,
            params: {
              num_ctx: 32768,
              temperature: 0.7,
              top_p: 0.9,
              thinking: false,
            },
          }
        ]
      }
    }
  }
}

Per-model agents.defaults.models["ollama/<model>"].params.num_ctx also works; the explicit provider model entry wins if both are set.

Thinking control

Native local Ollama compaction summaries default to thinking off. This keeps summarization from using its default three-minute request window for reasoning; Qwen3.5 treats low as thinking enabled rather than a reduced thinking budget. An explicit agents.defaults.compaction.thinkingLevel overrides this preference. Existing per-model params.think/params.thinking settings keep their normal precedence. Hosted routes keep their compaction defaults.

OpenClaw forwards thinking as Ollama expects it: top-level think, not options.think. Auto-discovered models whose /api/show reports a thinking capability expose /think low, /think medium, /think high, and /think max; non-thinking models expose only /think off. When /api/show also reports thinking.values, discovery caches a model-specific mapping. Boolean models send false or true; graded models send their supported effort strings. Advertised xhigh is also available through /think xhigh and --thinking xhigh.

Existing selections map to the nearest supported tier using OpenClaw's shared thinking ladder. For example, a model advertising low, medium, and xhigh receives xhigh for High and Maximum. If the model cannot disable thinking, Off uses its lowest supported tier. OpenClaw keeps its existing Off default rather than adopting thinking.default. Discovery caches these mappings with model metadata; turns do not fetch them again.

When replaying an assistant message, native requests retain its available reasoning in Ollama's separate thinking field alongside text and tool calls. This lets tool follow-ups reuse reasoning retained by the session's history policy without mixing it into visible answer text.

openclaw agent --model ollama/gemma4 --thinking off
openclaw agent --model ollama/gemma4 --thinking low

Or set a model default:

{
  agents: {
    defaults: {
      models: {
        "ollama/gemma4": {
          params: { thinking: "low" },
        },
      },
    },
  },
}

Per-model params.think/params.thinking can disable or force API thinking for a specific model. OpenClaw preserves that explicit config when the active run only has the implicit off default; a non-off runtime command such as /think medium still overrides it. A model marked reasoning: false suppresses enabled runtime selections. However, a mandatory-thinking model still uses its lowest supported tier for Off or a configured false, keeping reasoning out of the answer even when its visibility is disabled.

Reasoning models

Models named deepseek-r1, reasoning, reason, or think are treated as reasoning-capable by default — no extra config needed:

ollama pull deepseek-r1:32b
Model costs

Ollama runs locally and is free, so all model costs are 0 for both auto-discovered and manually defined models.

Memory embeddings

The bundled Ollama plugin registers a memory embedding provider for memory search. It uses the configured Ollama base URL and API key, calls /api/embed, and batches multiple memory chunks into one input request when possible.

When proxy.enabled=true, embedding requests to the exact host-local loopback origin derived from the configured baseUrl use OpenClaw's guarded direct path instead of the managed forward proxy. The configured hostname must itself be localhost or a loopback IP literal — DNS names that merely resolve to loopback still use the managed proxy path. LAN, tailnet, private-network, and public Ollama hosts always stay on the managed proxy path, and redirects to another host/port do not inherit trust. proxy.loopbackMode: "proxy" routes loopback traffic through the proxy anyway; proxy.loopbackMode: "block" denies it before connecting — see Managed proxy.

PropertyValue
Default modelnomic-embed-text
Auto-pullNo; pull the model on the Ollama host first
Embedding concurrencyProvider-owned; no memory-search tuning key is required

Before indexing memory, run ollama pull nomic-embed-text on the configured Ollama host (or pull the model selected by memory.search.model). A missing model returns HTTP 404; OpenClaw does not download it automatically.

Query-time embeddings use retrieval prefixes for models that require or recommend them: nomic-embed-text, qwen3-embedding, and mxbai-embed-large. Document batches stay raw, so existing indexes need no format migration.

Embedding concurrency and batching behavior are owned by the Ollama memory provider. For a remote embedding host, use the supported remote.baseUrl and remote.apiKey fields to keep auth scoped to that host:

{
  memory: {
    search: {
      provider: "ollama",
      model: "nomic-embed-text",
      remote: {
        baseUrl: "http://gpu-box.local:11434",
        apiKey: "ollama-local",
      },
    },
  },
}
Streaming configuration

Ollama uses the native API (/api/chat) by default, which supports streaming and tool calling together — no special config needed.

Like Chat Completions, native Ollama withholds durable reply blocks until the text phase is known at message completion, even with blockStreamingBreak: "text_end". Tool narration stays out of replies; ordinary finals and length-limited answers remain deliverable. Independent live previews can still update when enabled. See Pending text phases.

For native requests, /think off, openclaw agent --thinking off, and plugin api.runtime.llm.complete({ reasoning: "off" }) calls send top-level think: false unless an explicit params.think/params.thinking is configured or discovery reports a model that cannot disable thinking. Direct completions that omit reasoning keep the model default. Without a discovered thinking descriptor, /think low|medium|high send the matching effort string. Verified full-effort Ollama Cloud families such as GLM 5.2, GLM 5.3, GLM 5.3 Flash, Kimi K3, DeepSeek V4, and DeepSeek V4.1 Flash also send native think: "max" for /think max; other models and local servers keep the compatible think: "high" mapping. Native max applies to the ollama-cloud provider and to any ollama provider whose base URL is https://ollama.com. A :cloud model reached through a local Ollama server keeps high, because Ollama 0.21.2 and earlier reject max.

GLM 5.3 and GLM 5.3 Flash on Ollama Cloud cannot turn thinking off: their /api/show thinking values have no false, and think: false makes them answer with their reasoning inline. For these models, /think off and a configured false send their lowest level, think: "low", instead, in agent turns and in one-shot completions such as openclaw infer model run.

For the OpenAI-compatible endpoint instead, see "Legacy OpenAI-compatible mode" above — streaming and tool calling may not work together there.