# Ollama troubleshooting

> Fixes for common Ollama failures in OpenClaw

- 网址：https://funcoding.ai/agents/openclaw/providers/ollama/troubleshooting/
- 来源：OpenClaw 官方文档原文（英文），MIT 许可，同步于 2026-10-11
- 官方原文：https://docs.openclaw.ai/zh-CN/providers/ollama/troubleshooting

---
## Troubleshooting

<details>
<summary>WSL2 crash loop (repeated reboots)</summary>

On WSL2 with NVIDIA/CUDA, the official Ollama Linux installer creates an
`ollama.service` systemd unit with `Restart=always`. If that service
autostarts and loads a GPU-backed model during WSL2 boot, Ollama can pin
host memory while loading; Hyper-V memory reclaim cannot always reclaim
those pages, so Windows can terminate the WSL2 VM, systemd restarts
Ollama, and the loop repeats.

Evidence: repeated WSL2 reboots/terminations, high CPU in `app.slice` or
`ollama.service` right after WSL2 startup, and SIGTERM from systemd rather
than the Linux OOM killer.

OpenClaw logs a startup warning when it detects WSL2, `ollama.service`
enabled with `Restart=always`, and visible CUDA markers.

Mitigation:

```bash
sudo systemctl disable ollama
```

On the Windows side, add this to `%USERPROFILE%\.wslconfig`, then run
`wsl --shutdown`:

```ini
[experimental]
autoMemoryReclaim=disabled
```

Or shorten keep-alive / start Ollama manually only when needed:

```bash
export OLLAMA_KEEP_ALIVE=5m
ollama serve
```

See [ollama/ollama#11317](https://github.com/ollama/ollama/issues/11317).

</details>

<details>
<summary>Ollama not detected</summary>

Confirm Ollama is running and is in the agent's model scope. For ambient
localhost discovery, set `OLLAMA_API_KEY` (or an auth profile). An explicit
self-hosted endpoint is discovered whether or not it lists models:

```bash
ollama serve
curl http://localhost:11434/api/tags
```

</details>

<details>
<summary>No models available</summary>

Pull the model locally, or define it explicitly in
`models.providers.ollama`:

```bash
ollama list  # See what's installed
ollama pull gemma4
ollama pull gpt-oss:20b
ollama pull llama3.3     # Or another model
```

</details>

<details>
<summary>Connection refused</summary>

```bash
# Check if Ollama is running
ps aux | grep ollama

# Or restart Ollama
ollama serve
```

</details>

<details>
<summary>Remote host works with curl but not OpenClaw</summary>

Verify from the same machine and runtime that runs the Gateway:

```bash
openclaw gateway status --deep
curl http://ollama-host:11434/api/tags
```

Common causes:

- `baseUrl` points at `localhost`, but the Gateway runs in Docker or on another host.
- The URL uses `/v1`, selecting OpenAI-compatible behavior instead of native Ollama.
- The remote host needs firewall or LAN binding changes.
- The model is on your laptop's daemon but not the remote one.

</details>

<details>
<summary>Model outputs tool JSON as text</summary>

Usually the provider is in OpenAI-compatible mode, or the model cannot
handle tool schemas. Prefer native mode:

```json5
{
  models: {
    providers: {
      ollama: {
        baseUrl: "http://ollama-host:11434",
        api: "ollama",
      },
    },
  },
}
```

If a small local model still fails on tool schemas, set
`compat.supportsTools: false` on that model entry and retest.

</details>

<details>
<summary>Repeated tool errors stop the turn</summary>

OpenClaw stops after three consecutive identical failures for the same tool
and arguments, including repeated unknown tool IDs. This protection is always
active; enabling `tools.loopDetection` is not required.

On Ollama 0.40.1, some model templates can cause the server to mistake the
`<tool_call>` marker for Tool Search's `tool_call` function. OpenClaw avoids
this collision with transport-only tool aliases, so Tool Search can remain
enabled. Execution and session history keep the original tool names; raw
provider requests may show names such as `openclaw_tool_call`.

This translation covers native Ollama and the Ollama plugin's identified
OpenAI-compatible chat-completions route. Other providers are unchanged.
Custom Ollama templates with different markers may need separate diagnosis.

Check the arguments in the recorded error. If the model repeatedly invents
tool names or cannot use the exposed schemas, switch to a model with native
tool calling and start a new turn. Changed errors and successful retries
reset the count. See [Tool-loop detection](https://funcoding.ai/agents/openclaw/tools/loop-detection/).

</details>

<details>
<summary>Kimi or GLM returns garbled symbols</summary>

Hosted Kimi/GLM responses that are long, non-linguistic symbol runs are
treated as a failed provider call rather than a successful reply, so
normal retry/fallback/error handling takes over instead of persisting
corrupted text into the session.

If it recurs, capture the model name, the current session file, and
whether the run used `Cloud + Local` or Ollama Cloud, then try a fresh
session and a fallback model:

```bash
openclaw infer model run --model ollama/kimi-k2.5:cloud --prompt "Reply with exactly: ok" --json
openclaw models set ollama/gemma4
```

</details>

<details>
<summary>Cold local model times out</summary>

Large local models can need a long first load. Scope the timeout to the
Ollama provider and optionally keep the model loaded between turns:

```json5
{
  models: {
    providers: {
      ollama: {
        timeoutSeconds: 300,
        models: [
          {
            id: "gemma4:26b",
            name: "gemma4:26b",
            params: { keep_alive: "15m" },
          },
        ],
      },
    },
  },
}
```

If the host itself is slow to accept connections, `timeoutSeconds` also
extends the guarded connect timeout for this provider.

</details>

<details>
<summary>Large-context model is too slow or runs out of memory</summary>

Many models advertise contexts larger than your hardware can run
comfortably. Native requests forward the effective `contextTokens` unless
`params.num_ctx` overrides it. Cap both OpenClaw's budget and Ollama's request
context for predictable first-token latency:

```json5
{
  models: {
    providers: {
      ollama: {
        maxTokens: 8192,
        models: [
          {
            id: "qwen3.5:9b",
            name: "qwen3.5:9b",
            contextTokens: 32768,
            params: { num_ctx: 32768, thinking: false },
          },
        ],
      },
    },
  },
}
```

Lower the model entry's `contextTokens` if OpenClaw sends too much prompt. Lower
`params.num_ctx` if Ollama's runtime context is too large for the machine.
Lower `maxTokens` if generation runs too long.

</details>

<div class="callout callout-note">

More help: [Troubleshooting](https://funcoding.ai/agents/openclaw/help/troubleshooting/) and [FAQ](https://funcoding.ai/agents/openclaw/help/faq/).

</div>
