# Gateway service and process

> The managed gateway service not running, macOS launchd faults, and high-memory exits

- 网址：https://funcoding.ai/agents/openclaw/gateway/troubleshooting/gateway-service-and-process/
- 来源：OpenClaw 官方文档原文（英文），MIT 许可，同步于 2026-10-11
- 官方原文：https://docs.openclaw.ai/zh-CN/gateway/troubleshooting/gateway-service-and-process

---
## Gateway service not running

Use when the service is installed but the process does not stay up.

```bash
openclaw gateway status
openclaw status
openclaw logs --follow
openclaw doctor
openclaw gateway status --deep   # also scan system-level services
```

Look for:

- `Runtime: stopped` with exit hints.
- Service config mismatch (`Config (cli)` vs `Config (service)`).
- Port/listener conflicts.
- Extra launchd/systemd/schtasks installs when `--deep` is used.
- `Other gateway-like services detected (best effort)` cleanup hints.

<details>
<summary>Common signatures</summary>

- `Gateway start blocked: set gateway.mode=local` or `existing config is missing gateway.mode` → local gateway mode is not enabled, or the config file was clobbered and lost `gateway.mode`. Fix: set `gateway.mode="local"` in your config, or re-run `openclaw onboard --mode local` / `openclaw setup` to restamp the expected local-mode config. If you are running OpenClaw via Podman, the default config path is `~/.openclaw/openclaw.json`.
- `refusing to bind gateway ... without auth` → non-loopback bind without a valid gateway auth path (token/password, or trusted-proxy where configured).
- `another gateway instance is already listening` / `EADDRINUSE` → port conflict.
- `Other gateway-like services detected (best effort)` → stale or parallel launchd/systemd/schtasks units exist. Most setups should keep one gateway per machine; if you do need more than one, isolate ports + config/state/workspace. See [/gateway#multiple-gateways-same-host](https://funcoding.ai/agents/openclaw/gateway/#multiple-gateways-same-host).
- `System-level OpenClaw gateway service detected` from doctor → a systemd system unit exists while the user-level service is missing. Remove or disable the duplicate before allowing doctor to install a user service, or set `OPENCLAW_SERVICE_REPAIR_POLICY=external` if the system unit is the intended supervisor.
- `Gateway service port does not match current gateway config` → the installed supervisor still pins the old `--port`. Run `openclaw doctor --fix` or `openclaw gateway install --force`, then restart the gateway service.

</details>

Related:

- [Background exec and process tool](https://funcoding.ai/agents/openclaw/gateway/background-process/)
- [Configuration](https://funcoding.ai/agents/openclaw/gateway/configuration/)
- [Doctor](https://funcoding.ai/agents/openclaw/gateway/doctor/)

## macOS gateway silently stops responding, then resumes when you touch the dashboard

Use when channels (Telegram, WhatsApp, etc.) on a macOS host go quiet for minutes to hours at a time, and the gateway appears to come back the moment you open the Control UI, SSH in, or otherwise interact with the host. There is usually no obvious symptom in `openclaw status` because by the time you look the gateway is alive again.

```bash
ls ~/.openclaw/logs/stability/ | tail -5
openclaw gateway stability --bundle latest
pmset -g log | grep -iE "sleep|wake|maintenance" | tail -50
launchctl print gui/$UID/ai.openclaw.gateway | grep -E "state|last exit|runs"
```

Look for:

- One or more `*-uncaught_exception.json` bundles in `~/.openclaw/logs/stability/` with `error.code` set to a transient network code such as `ENETDOWN`, `ENETUNREACH`, `EHOSTUNREACH`, or `ECONNREFUSED`.
- `pmset -g log` lines like `Entering Sleep state due to 'Maintenance Sleep'` or `en0 driver is slow (msg: WillChangeState to 0)` aligned with the crash timestamps. Power Nap / Maintenance Sleep briefly puts the Wi-Fi driver into state 0; any outbound `connect()` that lands in that window can fail with `ENETDOWN` even on a host that otherwise has full network connectivity.
- `launchctl print` output showing `state = not running` with multiple recent `runs` and an exit code, especially when the gap between crash and the next launch is on the order of an hour rather than seconds. macOS launchd applies an undocumented respawn-protection gate after a crash burst that can stop honoring `KeepAlive=true` until an external trigger such as interactive login, dashboard connection, or `launchctl kickstart` re-arms it.

Common signatures:

- A stability bundle whose `error.code` is `ENETDOWN` or a sibling code, with the call stack pointing into Node `net` `lookupAndConnect` / `Socket.connect`. OpenClaw `2026.5.26` and newer classify these as benign transient network errors so they no longer propagate to the top-level uncaught handler; if you are on an older release, upgrade first.
- Long quiet periods that end the instant you connect to the Control UI or SSH into the host: the user-visible activity is what re-arms launchd's respawn gate, not anything the dashboard does to the gateway.
- `runs` count incrementing across the day with no corresponding `received SIG*; shutting down` line in `~/Library/Logs/openclaw/gateway.log`: clean shutdowns log a signal; transient crashes do not.

What to do:

1. **Upgrade the gateway** if you are running a release before `2026.5.26`. After upgrading, future `ENETDOWN` errors are logged as warnings instead of terminating the process.
2. **Reduce maintenance sleep activity** on Mac mini / desktop hosts meant to run as always-on servers:

   ```bash
   sudo pmset -a sleep 0 disksleep 0 standby 0 powernap 0
   ```

   This significantly reduces, but does not entirely eliminate, the underlying driver flap. The system can still perform some maintenance sleeps for TCP keepalive and mDNS upkeep regardless of these flags.

3. **Add a liveness watchdog** so a future crash burst that gets parked by launchd is caught quickly:

   ```bash
   # Example launchd-aware liveness check, suitable for a 5-minute cron or LaunchAgent
   state=$(launchctl print gui/$UID/ai.openclaw.gateway 2>/dev/null | awk -F'= ' '/state =/ {print $2; exit}')
   if [ "$state" != "running" ]; then
     launchctl kickstart -k gui/$UID/ai.openclaw.gateway
   fi
   ```

   The point is to externally re-arm the respawn gate; `KeepAlive=true` alone is not sufficient on macOS after a crash burst.

Related:

- [macOS platform notes](https://funcoding.ai/agents/openclaw/platforms/macos/)
- [Logging](https://funcoding.ai/agents/openclaw/logging/)
- [Doctor](https://funcoding.ai/agents/openclaw/gateway/doctor/)

## macOS launchd supervisor loop with duplicate gateway/node LaunchAgents

Use this when a macOS install keeps restarting every few seconds, `openclaw`
health checks flap between healthy and unavailable, and channel dispatch stalls
even though the service appears to be running.

This happens when both `ai.openclaw.gateway` and
`ai.openclaw.node` LaunchAgents are active and each injects
`OPENCLAW_LAUNCHD_LABEL`. In that state OpenClaw can detect launchd
supervision, try to hand restart back to launchd, and fall into a fast
`EADDRINUSE`/respawn loop instead of one stable gateway process.

```bash
for i in 1 2 3 4; do
  ps aux | grep 'openclaw.*index.js' | grep -v grep | awk '{print $2}'
  sleep 10
done

openclaw gateway status --deep
openclaw node status
launchctl print gui/$UID/ai.openclaw.gateway | grep -E 'state|last exit|runs'
tail -n 80 ~/Library/Logs/openclaw/gateway.log
```

Look for:

- More than one gateway PID across the 30-second sample instead of one stable
  process.
- `EADDRINUSE`, `another gateway instance is already listening`, or repeated
  restart/handoff lines in `gateway.log`.
- Both `~/Library/LaunchAgents/ai.openclaw.gateway.plist` and
  `~/Library/LaunchAgents/ai.openclaw.node.plist` loaded at the same time on a
  host that should only run one managed gateway service.

What to do:

1. If this host should only run the Gateway service, remove the managed node
   service through OpenClaw. **Skip this step** if you actively rely on the node
   service for remote node features; uninstalling it stops those features on
   this host:

   ```bash
   openclaw node uninstall
   ```

2. Install a persistent Gateway wrapper that clears the inherited launchd
   markers before starting OpenClaw. Use the supported `--wrapper` option; do
   not edit the generated file under `~/.openclaw/service-env/`, because service
   reinstall, update, and doctor repair regenerate that file:

   ```bash
   mkdir -p ~/.local/bin
   cat >~/.local/bin/openclaw-launchd-workaround <<'EOF'
   #!/bin/sh
   set -eu
   unset OPENCLAW_LAUNCHD_LABEL LAUNCH_JOB_LABEL LAUNCH_JOB_NAME XPC_SERVICE_NAME || true
   exec openclaw "$@"
   EOF
   chmod 700 ~/.local/bin/openclaw-launchd-workaround

   openclaw gateway install \
     --wrapper ~/.local/bin/openclaw-launchd-workaround \
     --force
   ```

   `gateway install` persists the wrapper path across forced reinstalls,
   updates, and doctor repairs.

3. Verify that the Gateway is stable and serving RPC, not merely listening:

   ```bash
   openclaw gateway status --deep --require-rpc

   for i in 1 2 3 4; do
     ps aux | grep 'openclaw.*index.js' | grep -v grep | awk '{print $2}'
     sleep 10
   done
   ```

   The PID sample should show one stable process instead of a rotating set of
   PIDs, and inbound channel dispatch should resume.

4. After upgrading to a release where the underlying dual-LaunchAgent loop is
   fixed, remove the workaround and reinstall the normal managed service:

   ```bash
   OPENCLAW_WRAPPER= openclaw gateway install --force
   rm ~/.local/bin/openclaw-launchd-workaround
   ```

Related:

- [Gateway on macOS](https://funcoding.ai/agents/openclaw/platforms/mac/bundled-gateway/)
- [Doctor](https://funcoding.ai/agents/openclaw/gateway/doctor/)
- [Gateway CLI](https://funcoding.ai/agents/openclaw/cli/gateway/)

## Native aborts on Linux (SIGABRT)

`malloc(): invalid next->prev_inuse (unsorted)` followed by systemd
`status=6/ABRT` indicates detected native heap corruption. It does not identify
the corrupting code, and it is different from a kernel OOM kill. A JavaScript
signal handler or stability bundle cannot reliably capture a native `abort()`.
Arrange OS core capture **before** another failure.

Startup logs include `native runtime` (PID, platform, architecture, Node, V8,
libuv, OpenSSL and SQLite versions) and `worker startup state` (tracked Worker
count, starts/retirements by script, and shared compute admission counters).
These are startup facts, not a snapshot of the moment of failure. Retain them
with the crash timestamp, journal, exact OpenClaw build and Node executable.

For a systemd system service, substitute the installed unit name below. For a
user service, use `systemctl --user` without `sudo`; its hard core limit cannot
exceed the user manager's inherited limit.

```bash
sudo systemctl edit openclaw-gateway.service
# Add this drop-in:
# [Service]
# LimitCORE=infinity

sudo systemctl daemon-reload
sudo systemctl show openclaw-gateway.service -p LimitCORE -p MainPID
sysctl kernel.core_pattern
```

Apply the limit at the next operator-coordinated service restart. Then check
`/proc/<gateway-pid>/limits` for the running process's `Max core file size`;
changing the unit does not change an already running process's limit.

Choose the capture backend indicated by `kernel.core_pattern`:

- **systemd-coredump:** verify that the distribution's handler is installed and
  enabled. Inspect `coredump.conf` storage and size limits; a multi-gigabyte
  Gateway needs enough disk space and `ProcessSizeMax`/`ExternalSizeMax` to keep
  its core. After a crash, use `sudo coredumpctl info <crashed-pid>` and
  `sudo coredumpctl debug <crashed-pid>`.
- **Apport:** inspect `/var/log/apport.log` and `/var/crash`. An existing report
  for the same Node executable/user can suppress a later report. Archive that
  exact stale `.crash` report outside `/var/crash` in a root-only directory,
  then clear any matching stale sidecars according to the distribution's
  Apport procedure. Do not erase unrelated reports. Verify Apport is enabled
  and actually accepts the next report; `code=dumped` alone does not prove a
  core was saved.
- **Direct core files:** if the installed handler cannot retain this crash,
  an administrator can temporarily replace it with a private absolute path:

  ```bash
  # Record the old value so it can be restored after the investigation.
  sysctl kernel.core_pattern
  sudo install -d -m 0700 -o <gateway-user> -g <gateway-group> /var/lib/openclaw-cores
  sudo sysctl -w 'kernel.core_pattern=/var/lib/openclaw-cores/core.%e.%p.%t'
  ```

  `core_pattern` is host-wide: this replaces capture for other processes too.
  Ensure the directory is writable by the Gateway service user and accessible
  inside any service filesystem sandbox. Restore the previous pattern after
  capture. This temporary `sysctl` setting does not survive reboot.

Validate the chosen backend with a disposable process under equivalent service
limits and identity, never by aborting the live Gateway. Open a captured core
with the **matching** Node binary and debug symbols (`gdb /path/to/node
/path/to/core` for direct files), then run `thread apply all bt`. Preserve all
thread stacks: the aborting thread can be detecting damage caused elsewhere.
Core files contain process memory, including credentials and message content;
keep them private and share only reviewed, redacted evidence.

## Gateway exits during high memory use

Use when the Gateway disappears under load, the supervisor reports an OOM-style restart, or logs show `memory pressure: level=critical`.

```bash
openclaw gateway status --deep
openclaw logs --follow
openclaw gateway stability --bundle latest
openclaw gateway diagnostics export
```

Look for:

- `memory pressure: level=critical` with `reason=rss_threshold`, `heap_threshold`, or `rss_growth`.
- RSS, heap, threshold, and growth values in that log line.
- Existing stability bundles from fatal exits, shutdown timeouts, or restart startup failures, when available.

Common signatures:

- `memory pressure: level=critical` appears in gateway logs → OpenClaw detected critical memory pressure and recorded the available in-process memory facts.
- `reason=heap_threshold` → lower prompt/session pressure or reduce concurrent work first. For a managed service, compare the configured controls and install-time recommendation in `Gateway heap:` from `openclaw gateway status` with the runtime measurement. Reinstalling preserves existing stored heap settings; it does not automatically replace an older value with the current recommendation.
- `reason=rss_growth` → the RSS floor kept rising across consecutive sampling windows. Check the latest logs for a large import, runaway tool output, repeated retries, or a batch of queued agent work.
- Critical memory pressure appears in logs but no bundle exists → capture `openclaw gateway diagnostics export` after the event for the available operational evidence. Pressure events do not automatically write bundles.

An administrator can also [sample allocations](https://funcoding.ai/agents/openclaw/gateway/diagnostics/#sampling-heap-profile) with `openclaw gateway call diagnostics.heapProfile --timeout 30000` on Node or OpenClaw's Bun runtime. This captures current allocation activity, not a past spike or all native memory. Older bundles remain readable with `openclaw gateway stability --bundle latest`.

Review the sanitized diagnostics export before attaching it to a bug report; avoid copying raw logs.

Node's automatic heap ceiling can be roughly 4 GiB on a large host. That is a default sizing decision, not a general 64-bit address-space ceiling. `--max-old-space-size` controls V8 old space; the measured total V8 heap ceiling also includes other heap spaces. RSS additionally includes native allocations, buffers, and other process memory. A higher heap ceiling does not preallocate the ceiling, but it still needs enough real capacity and headroom under sustained load.

For a foreground Node Gateway, set a native heap flag before Node starts, for example on a host with sufficient capacity:

```bash
NODE_OPTIONS="--max-old-space-size=16384" openclaw gateway run
```

For a custom supervisor or Docker runtime command, place `--max-old-space-size=16384` immediately after `node`, before the OpenClaw entry script, or set `NODE_OPTIONS` in that process or container's launch environment. Docker image build-time heap options do not configure the runtime Gateway. An OpenClaw config or dotenv value loaded after Node starts cannot resize its heap. `NODE_OPTIONS` can also reach spawned Node children, so prefer a direct Node argument when only the Gateway should receive the budget.

For managed Node services, use the [managed Gateway heap policy](https://funcoding.ai/agents/openclaw/cli/gateway/#manage-the-gateway-service) and inspect both managed launch arguments and operator-owned environment overrides before changing them. Native argv overrides the same option in `NODE_OPTIONS`; percentage old-space sizing takes precedence over absolute old-space sizing. Regeneration preserves stored argv but does not add an automatic heap flag when an operator override owns `NODE_OPTIONS`. Installer-shell `NODE_OPTIONS` does not become a service override. Runtime pressure diagnostics use the effective V8 heap ceiling and physical/reported constraint headroom; an oversized explicit heap setting does not raise the RSS alert threshold above physical capacity. Pressure warnings are diagnostic evidence, not heap limits or automatic restart triggers.

RSS growth detection compares minimum RSS values from completed five-minute windows and requires two consecutive increases. Growth accumulates while these floors keep rising; a flat or falling floor, a sampling gap over ten minutes, or a clock rollback resets the trend. This filters ordinary GC peaks while detecting smaller sustained increases. The logged `rssGrowth` and `windowMs` describe the accumulated floor increase and elapsed time between those minima, which can exceed ten minutes.

On Node, growth warnings and critical events use 4% and 8% of the smaller of the measured V8 heap ceiling and available process capacity, with minimum thresholds of 512 MiB and 1 GiB. Process capacity uses the reported constraint bounded by physical RAM, or physical RAM when no constraint is reported. With a 16 GiB heap and sufficient RAM, those thresholds are about 655 MiB and 1.28 GiB. Unknown heap limits and Bun retain the 512 MiB/1 GiB growth thresholds; Bun's existing absolute memory caps are unchanged. Absolute RSS and heap pressure checks still run on every sample.

Related:

- [Gateway health](https://funcoding.ai/agents/openclaw/gateway/health/)
- [Diagnostics export](https://funcoding.ai/agents/openclaw/gateway/diagnostics/)
- [Sessions](https://funcoding.ai/agents/openclaw/cli/sessions/)
