Ollama advanced configuration
Context windows, thinking control, costs, embeddings, and streaming for Ollama
Advanced configuration
Native Ollama requests keep exact session identities and runtime facts in the first user message, after the shared system and tool prefix. This allows local prompt caches to reuse that prefix across equivalent subagent spawns while preserving earlier message bytes on follow-up turns. Cache reuse still depends on the model template, available cache slots, and unchanged instructions/tools.
Legacy OpenAI-compatible mode
Tool calling is not reliable in this mode. Use it only when a proxy needs OpenAI format and you do not depend on native tool calling.
Set api: "openai-completions" explicitly for a proxy behind
/v1/chat/completions:
{
models: {
providers: {
ollama: {
baseUrl: "http://ollama-host:11434/v1",
api: "openai-completions",
injectNumCtxForOpenAICompat: true, // default: true
apiKey: "ollama-local",
models: [...]
}
}
}
}This mode may not support streaming and tool calling simultaneously; you
may need params: { streaming: false } on the model in
models.providers.ollama.models. You can also set it under
agents.defaults.models["ollama/<model>"].params. For openai-completions,
this sends stream: false upstream and delivers the completed reply with
its usage, reasoning, and tool calls. Streaming remains enabled when unset;
this setting does not change native Ollama /api/chat requests.
OpenClaw injects options.num_ctx by default in this mode so Ollama does
not silently fall back to a 4096-token context. If your proxy rejects
unknown options fields, disable it:
{
models: {
providers: {
ollama: {
baseUrl: "http://ollama-host:11434/v1",
api: "openai-completions",
injectNumCtxForOpenAICompat: false,
apiKey: "ollama-local",
models: [...]
}
}
}
}Context windows
For auto-discovered models, OpenClaw uses the context window /api/show
reports, including larger PARAMETER num_ctx values from custom
Modelfiles; otherwise it falls back to OpenClaw's default Ollama context
window.
Per-model contextWindow declares native window metadata, and per-model
contextTokens caps active input. Provider-level maxTokens remains an
output-token default; a model entry can override it. Native
/api/chat requests set options.num_ctx from a positive params.num_ctx
first, then from the effective model contextTokens when present. Local
discovery normally caps contextTokens at 32,768 (or the model's smaller
native window), so OpenClaw can override a smaller Modelfile context even
without an explicit params.num_ctx. Invalid, zero, negative, or non-finite
params.num_ctx values are ignored. Only when neither value is available
does Ollama choose its own model, Modelfile, OLLAMA_CONTEXT_LENGTH, or
VRAM-based default; the native adapter does not fall back directly to the
advertised contextWindow. After upgrading an older configuration, run
openclaw doctor --fix. Doctor preserves current contextTokens caps without
creating a stronger model or provider num_ctx pin; uncapped legacy native
entries still migrate their older context budgets. Existing explicit
params.num_ctx values remain authoritative, including pins an older Doctor
already wrote. Review or remove an oversized existing pin to let
contextTokens drive the request again. Use params.num_ctx to override
the native request context explicitly. The
OpenAI-compatible adapter still injects options.num_ctx by default from
params.num_ctx, then the matching model entry's contextTokens or
contextWindow; disable with
injectNumCtxForOpenAICompat: false if the upstream rejects options.
Native model entries also accept common Ollama runtime options under
params, forwarded as native /api/chat options: num_keep, seed,
num_predict, top_k, top_p, min_p, typical_p, repeat_last_n,
temperature, repeat_penalty, presence_penalty, frequency_penalty,
stop, num_batch, num_gpu, main_gpu, use_mmap, and num_thread.
Runtime sampling controls (temperature, topP, frequencyPenalty,
presencePenalty, and seed) override the matching model defaults,
including explicit zero values. The Gateway's Chat Completions API maps
top_p, frequency_penalty, and presence_penalty to these controls.
With temperature: 0, OpenClaw still normalizes top_p to 1 for greedy
sampling after applying overrides.
A few keys (format, keep_alive, truncate, shift) are forwarded as
top-level request fields instead of nested options. Local native chat
requests default to truncate: false and shift: false, so supporting
servers reject overflowing input instead of silently dropping history.
OpenClaw then attempts compaction and retries, or reports the failure.
Generation that fills the window can still produce a labeled partial reply.
This behavior is verified with Ollama 0.33.3; older servers may ignore the
fields. Explicit per-model values override these defaults. Hosted models
and the OpenAI-compatible endpoint keep their existing behavior.
OpenClaw only
forwards these Ollama request keys, so runtime-only params such as
streaming are never sent to Ollama. Use params.think (or
params.thinking) to set top-level think; false disables API-level
thinking for Qwen-style thinking models.
{
models: {
providers: {
ollama: {
models: [
{
id: "llama3.3",
contextWindow: 131072,
contextTokens: 32768,
maxTokens: 65536,
params: {
num_ctx: 32768,
temperature: 0.7,
top_p: 0.9,
thinking: false,
},
}
]
}
}
}
}Per-model agents.defaults.models["ollama/<model>"].params.num_ctx also
works; the explicit provider model entry wins if both are set.
Thinking control
Native local Ollama compaction summaries default to thinking off. This keeps
summarization from using its default three-minute request window for reasoning;
Qwen3.5 treats low as thinking enabled rather than a reduced thinking budget.
An explicit agents.defaults.compaction.thinkingLevel overrides this
preference. Existing per-model params.think/params.thinking settings
keep their normal precedence. Hosted routes keep their compaction defaults.
OpenClaw forwards thinking as Ollama expects it: top-level think, not
options.think. Auto-discovered models whose /api/show reports a
thinking capability expose /think low, /think medium, /think high,
and /think max; non-thinking models expose only /think off. When
/api/show also reports thinking.values, discovery caches a model-specific
mapping. Boolean models send false or true; graded models send their
supported effort strings. Advertised xhigh is also available through
/think xhigh and --thinking xhigh.
Existing selections map to the nearest supported tier using OpenClaw's
shared thinking ladder. For example, a model advertising low, medium,
and xhigh receives xhigh for High and Maximum. If the model cannot
disable thinking, Off uses its lowest supported tier. OpenClaw keeps its
existing Off default rather than adopting thinking.default. Discovery
caches these mappings with model metadata; turns do not fetch them again.
When replaying an assistant message, native requests retain its available
reasoning in Ollama's separate thinking field alongside text and tool
calls. This lets tool follow-ups reuse reasoning retained by the session's
history policy without mixing it into visible answer text.
openclaw agent --model ollama/gemma4 --thinking off
openclaw agent --model ollama/gemma4 --thinking lowOr set a model default:
{
agents: {
defaults: {
models: {
"ollama/gemma4": {
params: { thinking: "low" },
},
},
},
},
}Per-model params.think/params.thinking can disable or force API
thinking for a specific model. OpenClaw preserves that explicit config
when the active run only has the implicit off default; a non-off
runtime command such as /think medium still overrides it. A model marked
reasoning: false suppresses enabled runtime selections. However, a
mandatory-thinking model still uses its lowest supported tier for Off or a
configured false, keeping reasoning out of the answer even when its
visibility is disabled.
Reasoning models
Models named deepseek-r1, reasoning, reason, or think are treated
as reasoning-capable by default — no extra config needed:
ollama pull deepseek-r1:32bModel costs
Ollama runs locally and is free, so all model costs are 0 for both
auto-discovered and manually defined models.
Memory embeddings
The bundled Ollama plugin registers a memory embedding provider for
memory search. It uses the configured Ollama base URL
and API key, calls /api/embed, and batches multiple memory chunks into
one input request when possible.
When proxy.enabled=true, embedding requests to the exact host-local
loopback origin derived from the configured baseUrl use OpenClaw's
guarded direct path instead of the managed forward proxy. The configured
hostname must itself be localhost or a loopback IP literal — DNS names
that merely resolve to loopback still use the managed proxy path. LAN,
tailnet, private-network, and public Ollama hosts always stay on the
managed proxy path, and redirects to another host/port do not inherit
trust. proxy.loopbackMode: "proxy" routes loopback traffic through the
proxy anyway; proxy.loopbackMode: "block" denies it before connecting —
see Managed proxy.
| Property | Value |
|---|---|
| Default model | nomic-embed-text |
| Auto-pull | No; pull the model on the Ollama host first |
| Embedding concurrency | Provider-owned; no memory-search tuning key is required |
Before indexing memory, run ollama pull nomic-embed-text on the configured
Ollama host (or pull the model selected by memory.search.model). A missing
model returns HTTP 404; OpenClaw does not download it automatically.
Query-time embeddings use retrieval prefixes for models that require or
recommend them: nomic-embed-text, qwen3-embedding, and
mxbai-embed-large. Document batches stay raw, so existing indexes need
no format migration.
Embedding concurrency and batching behavior are owned by the Ollama
memory provider. For a remote embedding host, use the supported
remote.baseUrl and remote.apiKey fields to keep auth scoped to that
host:
{
memory: {
search: {
provider: "ollama",
model: "nomic-embed-text",
remote: {
baseUrl: "http://gpu-box.local:11434",
apiKey: "ollama-local",
},
},
},
}Streaming configuration
Ollama uses the native API (/api/chat) by default, which supports
streaming and tool calling together — no special config needed.
Like Chat Completions, native Ollama withholds durable reply blocks until
the text phase is known at message completion, even with
blockStreamingBreak: "text_end". Tool narration stays out of replies;
ordinary finals and length-limited answers remain deliverable. Independent
live previews can still update when enabled.
See Pending text phases.
For native requests, /think off, openclaw agent --thinking off, and
plugin api.runtime.llm.complete({ reasoning: "off" }) calls send top-level
think: false unless an explicit params.think/params.thinking is
configured or discovery reports a model that cannot disable thinking.
Direct completions that omit reasoning keep the model default.
Without a discovered thinking descriptor, /think low|medium|high send
the matching effort string. Verified full-effort
Ollama Cloud families such as GLM 5.2, GLM 5.3, GLM 5.3 Flash, Kimi K3,
DeepSeek V4, and DeepSeek V4.1 Flash also send native
think: "max" for /think max; other models and local servers keep the
compatible think: "high" mapping.
Native max applies to the ollama-cloud provider and to any ollama
provider whose base URL is https://ollama.com. A :cloud model reached
through a local Ollama server keeps high, because Ollama 0.21.2 and
earlier reject max.
GLM 5.3 and GLM 5.3 Flash on Ollama Cloud cannot turn thinking off: their
/api/show thinking values have no false, and think: false makes them
answer with their reasoning inline. For these models, /think off and a
configured false send their lowest level, think: "low", instead, in
agent turns and in one-shot completions such as openclaw infer model run.
For the OpenAI-compatible endpoint instead, see "Legacy OpenAI-compatible mode" above — streaming and tool calling may not work together there.