跳到正文
FunCoding

搜索

搜索文档、文章、Skill 和 MCP

Restart and supervision

openclaw gateway restart, install identity, external process supervisors, and Gateway profiling

Restart flags, which install owns the host service, external supervisors, and profiling. Part of the openclaw gateway reference.

Restart the Gateway

openclaw gateway restart
openclaw gateway restart --safe
openclaw gateway restart --safe --skip-deferral
openclaw gateway restart --force
openclaw gateway restart --wait 30s

Manual restart signals now use SIGUSR2. SIGUSR1 starts Node's inspector and no longer restarts the Gateway. Update scripts that send the old signal; prefer openclaw gateway restart for service-aware restarts.

If restart cannot verify a live serving owner, it leaves the process untouched. Run openclaw gateway status --deep, fix the reported startup failure (for example, a stopped Tailscale backend when Serve is configured), then run openclaw gateway start to wait for readiness. A loaded service or a running PID alone does not prove that the Gateway is serving. Reinstallation is not a remedy for an unresolved startup dependency or unknown process ownership.

--safe asks the running Gateway to preflight active work and schedule one coalesced restart after that work drains. The wait is bounded to 5 minutes; when the budget expires the restart is forced. --safe cannot combine with --force or --wait.

--skip-deferral bypasses only the safe-restart active-work deferral gate. It can move the Gateway into shutdown even while active-work blockers are reported, but the close-stage pending-reply drain still applies before the process exits. It requires --safe — use it when a deferral is stuck on a runaway task and reply delivery can still be allowed to settle.

--wait <duration> overrides the drain budget for a plain (non-safe) restart. Accepts bare milliseconds or unit suffixes ms, s, m, h, d (e.g. 30s, 5m, 1h30m); --wait 0 waits indefinitely. Not compatible with --force or --safe.

Native service stop deadlines still apply: the Gateway caps the drain at 315 seconds for systemd's 330-second limit, and 5 seconds for launchd's 20-second limit. Both leave time for cancellation and cleanup. These caps also apply to --wait 0. Longer model or heartbeat timeouts do not extend it. When available, the drain log reports the largest observed model request timeout for context.

Queued heartbeat wakes settle as gateway-draining when shutdown closes admission, including wakes waiting to retry. They cannot start another turn in the draining runtime. Already-running wakes and pending final reply writes retain their drain grace.

If work still ignores cancellation at the shutdown deadline under systemd or launchd, a native service stop or supervisor-owned restart logs the remaining work categories and pending owners (including command lanes and request origins), writes a diagnostic stability bundle, and exits with status 0. It does not reuse that unfinished runtime for an in-process restart. This lets a requested stop finish cleanly and lets the service manager start a fresh Gateway for a restart.

An explicit server-close failure retains exit status 1, including when final provider cleanup crosses the native shutdown deadline.

Admission-close logs name the shutdown trigger, for example stop (SIGTERM) or restart (SIGUSR2: config reload: gateway.bind). A signal alone does not identify its sender: Node does not expose the sender PID or command. Three occurrences of the same signal within five minutes in one process, or three recorded plain SIGTERM/SIGINT stops across process lifetimes, produce a hint to check openclaw gateway status --deep for another supervisor. Deep local status shows the last recorded shutdown reason and time for the selected state directory; on Linux, it also reports competing user and system service units when inspecting the native service. Failure outcomes keep their specific failure reason.

After a downgrade, the Gateway refuses databases whose schema is newer than the running build supports. The startup error and openclaw gateway status --deep report the found and supported schema versions, the writer build when recorded, and the refusing build. Run a build at least as new as the writer that supports those schemas, or stop the service and restore your pre-upgrade backup. Startup retains exit status 78 and parks a managed LaunchAgent when possible. A refused shared-state database cannot record a new lifecycle row; the error log explains the refusal, and deep status reports it instead of an unavailable shutdown record.

When the shared-state database cannot be read at all (for example, the file is damaged), openclaw gateway status --deep fails with exit status 1 and names the database path and read error instead of reporting a config read failure. Stop OpenClaw processes, then restore that file from a verified backup, as openclaw doctor also advises.

Foreground/manual Gateways, in-process restarts selected by OPENCLAW_NO_RESPAWN=1, and other supervisors retain exit status 1 when cleanup cannot finish before the shutdown deadline.

--force begins restart immediately and closes new admissions. The current CLI supplies the normal drain budget, capped by the native service shutdown deadline. Only work remaining at that deadline is canceled before cleanup. A safe restart whose deferral budget has already expired does not get a second drain budget. Plain restart normally uses the service-manager restart path.

Forced requests from older callers that supply no drain budget get at most 45 seconds to drain. Their 60-second replacement window reserves 10 seconds for cleanup and 5 seconds for replacement. This applies to interactive and update callers alike. If work remains when the drain expires, the Gateway records a warning with the remaining work categories in its restart history and log, then cancels that work through normal terminal recovery. Callers that supply a budget retain that budget, subject to native service deadlines.

During an upgrade, restart records its reason and drain options in the existing Gateway state without starting a schema migration while the old Gateway is still running. If no state database exists, it logs that intent recording was skipped and continues the restart.

When an updater invokes the installed gateway restart command, its existing update marker enables the five-minute startup watchdog after the managed process is observed running. This lets an older updater complete a slow first-hop startup without passing a new option. The watchdog includes migration, listener, and health phases; phase changes cannot extend its cap. Explicit readiness budgets supplied by newer update callers take precedence. Ordinary standalone restarts wait beyond the standard readiness budget only while the same running Gateway advances startup phases or acquires, renews, or completes an observed same-process migration, up to five minutes, then report still-starting (exit 2) with the last phase and openclaw gateway status --deep as the next step; startup without progress still fails at the standard budget (exit 1), while a newly observed migration lease gets one heartbeat interval plus polling grace before it is considered stalled, and its observed completion earns one fresh readiness window within the same cap. See Restart recovery.

On Windows, managed gateway start and gateway restart allow up to 90 minutes for cold startup, using three default update-step budgets for activation, loading, and readiness. Managed update restoration uses the same allowance unless an explicit update --timeout supplies its per-step budget. Implicit Windows update readiness checks use ten times the observed startup duration, bounded between 90 and 120 minutes; the ceiling cannot truncate the cold-start floor. This accommodates large agent databases on slow storage; a service that exits or fails a health check can still report failure earlier. A live Gateway that remains in startup at the deadline is left running and reported as still-starting; updates retain the readiness warning and recovery backups.

On Windows, a plain restart launched from a Gateway service process, including an agent's shell command, automatically uses the safe restart path. The running Gateway owns the deferred Scheduled Task handoff, so stopping its process tree cannot kill the caller before relaunch. This requires a reachable Gateway; the command acknowledges the restart request, not successor health. Use openclaw gateway status afterward to verify recovery.

The Windows handoff waits up to three minutes for the outgoing Gateway to exit, then requests a task launch. It records restart finished in logs/gateway-restart.log only after a different process with the expected executable and Gateway entrypoint listens on the configured port. This listener check uses the same 90-minute cold-start allowance; it does not prove channel readiness. A task marked Running or a successful launch request alone does not count as recovery.

If no replacement listener appears, the log records restart failed and a profile-aware openclaw gateway restart --force command to run from an external terminal. The handoff does not end its own Scheduled Task: on installations with Job Object containment, doing so could terminate the observer before it records the result. A stale running task can still require this external recovery.

On macOS, when openclaw gateway restart, stop, install, or uninstall runs inside the managed LaunchAgent's process tree, including an agent's shell command, OpenClaw detects that from launchd's service environment or, when a hand-written plist omits those variables, from process ancestry against the PID launchd reports for the job. Restart hands off to a detached helper so kickstart -k cannot kill the caller. Stop, install, and uninstall refuse and ask you to run the command from an external shell.

External terminals without Gateway-service markers, externally supervised Gateways, node services, and non-Windows callers keep their existing routing. Explicit --force, --wait, --preserve-definition, or --skip-deferral also retain their existing behavior and validation; they do not implicitly enable --safe.

Inline --password can be exposed in local process listings. Prefer --password-file, env, or a SecretRef-backed gateway.auth.password.

Install identity

Service management (install, start, stop, restart, uninstall, Doctor service repair, and self-update service handling) belongs to the install that owns the host service. That is the canonical .openclaw directory under the OS account home, or the .openclaw-<profile> directory a named profile projects there. Named profiles use distinct native service identities.

OPENCLAW_HOME may explicitly select the OS account home, including a filesystem alias of that home. Both the process home (HOME or USERPROFILE) and the effective OpenClaw home must resolve to the account home. An OPENCLAW_HOME, OPENCLAW_STATE_DIR, or OPENCLAW_CONFIG_PATH that points elsewhere is treated as isolated state and skipped. A relocated or copied state tree cannot adopt and rewrite the account's host service.

Doctor also validates the environment saved in the installed service. A canonical OPENCLAW_HOME there does not prevent openclaw doctor --fix from entering maintenance and importing legacy credentials. Doctor remains the migration owner: it verifies the imported credentials and archives the original bytes before normal runtime reads resume.

On macOS and Windows, native service-managed profile names must be lowercase. Runtime-only profiles may still use uppercase, but case-distinct names such as Main and main share paths on normal case-insensitive filesystems and cannot safely own separate native services. On macOS, the lowercase names gateway and node are also unavailable for native service management because their historical LaunchAgent labels collide with the default Gateway and node-host services.

Named profiles must also use the native service identity derived from OPENCLAW_PROFILE. Unset OPENCLAW_LAUNCHD_LABEL, OPENCLAW_SYSTEMD_UNIT, or OPENCLAW_WINDOWS_TASK_NAME before service management; custom identities remain available for the default profile or runtime-only/external-supervisor setups.

On Linux, discovery also recognizes legacy openclaw-<profile> unit names. A custom system unit can belong to the default installation when its OpenClaw launcher, service account, profile, and state/config paths identify that installation. Discovery uses systemd's effective command and environment, including drop-ins and environment files. If multiple custom units match or a wrapper makes their identity unclear, specify the intended unit with OPENCLAW_SYSTEMD_UNIT; OpenClaw does not choose the first unit carrying its marker.

Doctor offers duplicate user-unit cleanup only when both managers' loaded Gateway commands identify the selected account, profile, state/config paths, and matching port selection. It rechecks the units after confirmation. Different or unverifiable identities leave the user unit in place. Cleanup removes only the confirmed user unit, then reports any remaining matching user unit or unverifiable discovery; another unit requires its own inspection and confirmation on a later Doctor run.

On Linux, openclaw gateway install --force refuses a sealed systemd service definition, or one whose write authority cannot be verified, before changing configuration, authentication tokens, or service files. The error keeps its SERVICE_DEFINITION_SEALED or SERVICE_DEFINITION_UNKNOWN prefix and adds a reason tag and next action, without printing private paths, config, environment values, or underlying inspection errors.

For [unsafe-permissions], inspect the named artifact category locally. The service directory is ~/.config/systemd/user; on a fresh install, its nearest existing ancestor may be ~/.config. The service state directory belongs to the selected profile. Check directory metadata, not file contents:

ls -ld ~/.config ~/.config/systemd ~/.config/systemd/user

Missing directories are normal on a fresh install. After confirming the affected path is yours and is not intentionally shared, remove group/other write access with chmod go-w <path> and retry the same command. Mode 0700 is appropriate for private directories. Do not recursively chmod, take ownership of system paths, or use sudo/--force to bypass the check. Foreign-owned files and sealed mounts require the deployment owner; inspection failures require restoring filesystem or native service-manager access first.

Type-wide service.d defaults are inspected as shared read-only inputs and do not require write access. Root-owned selected units and unit-specific drop-ins remain protected.

External supervisors

Set OPENCLAW_SUPERVISOR_MODE=external only when another process manager owns the Gateway lifecycle. In this mode:

  • openclaw gateway restart preserves the existing safe, forced, and bounded-wait behavior while targeting the verified running Gateway instead of launchd, systemd, or Task Scheduler. Exact-lock restart delivery runs inside that Gateway, so a replacement CLI does not migrate shared state before the old process hands off.
  • Native service install, start, stop, and uninstall operations are refused with guidance to use the external supervisor.
  • OpenClaw self-update is refused so the supervisor can stop the Gateway, replace and finalize the runtime, and restart it safely.
  • A fresh-process restart writes a bounded SQLite handoff before clean exit. If persistence fails, the Gateway falls back to an in-process restart instead of exiting without a consumable handoff.

An external supervisor can also claim durable ownership of shared-state writes:

OPENCLAW_SUPERVISOR_MODE=external \
  openclaw database ownership claim --manager gateway-supervisor --json

Before claiming, stop the Gateway through the external supervisor and stop any embedded agents. The CLI enforces this procedure: it refuses a live Gateway or embedded-agent owner and holds exclusive offline ownership while committing the claim. Also stop and verify every CLI, Doctor, updater, and native app process older than 2026.8.1 that can write the shared state database. Processes from before the ownership contract (#121069) do not understand the ownership row and cannot be retroactively fenced. Claim only after every remaining writer uses ownership-aware code and carries OPENCLAW_SUPERVISOR_MODE=external.

The claim is idempotent for the same stable manager identifier and refuses a different manager. There is no automatic claim or unclaim path. Once claimed, unmarked writable shared-state opens fail before permissions, schema migration, additive repair, compaction, or other mutation. Read-only access remains available. This is protection against accidental unmarked same-user writers, not an authentication or lease protocol.

For upgrades and rollbacks, have the supervisor create a consolidated WAL-consistent copied snapshot with no SQLite sidecars, then run the target release's own openclaw database preflight <copied-state.sqlite> --json before activation. Numeric schema versions alone do not prove that a same-version additive shape is compatible. See Database schemas.

OPENCLAW_SERVICE_REPAIR_POLICY=external remains a separate Doctor repair policy. It does not declare runtime ownership; supervisors that need both behaviors should set both variables.

External supervisors can negotiate and consume restart handoffs through the hidden machine contract:

openclaw gateway restart-handoff capabilities --json
openclaw gateway restart-handoff consume --expected-pid <pid> --json

Protocol version 1 supports the consume operation. Consumption validates the expected PID and bounded handoff fields inside one immediate SQLite transaction. An accepted handoff is deleted before success is returned, so concurrent or replayed consumers cannot both accept it. A PID mismatch is retained for the matching owner; missing, expired, and invalid rows do not authorize a restart.

Valid machine requests return JSON with exit code 0, including non-restart results. Invalid arguments return reason: "invalid-expected-pid" with exit code 2; state-store failures return reason: "store-unavailable" with exit code 1. Supervisors should check capabilities on the exact runtime or launcher they will use rather than infer support from an OpenClaw version string or read the private SQLite schema directly.

External supervisor implementations should also apply these acceptance rules:

  • Bound capability checks with a timeout that accounts for full CLI cold-start latency on the deployed runtime and storage, rather than assuming warm-start timing.
  • If capability negotiation or handoff consumption refuses replacement, exit promptly with a nonzero status so the process manager's recovery policy can run. Do not remain alive without a Gateway child or listener.
  • Treat supervisor process liveness as distinct from replacement startup and channel readiness. Report success only after the new Gateway owns its listener and /startupz returns status: "started"; monitor /readyz separately for configured-channel health, while /healthz proves liveness only.

Gateway profiling

  • OPENCLAW_GATEWAY_STARTUP_TRACE=1 logs phase timings during startup, including per-phase eventLoopMax delay and plugin lookup-table timings (installed-index, manifest registry, startup planning, owner-map work). The process.bootstrap breakdown includes earlier CLI imports, configuration, database admission, session inventory, and workspace readiness. Each bootstrap step records its process-relative start, duration, call count, and available fleet counts; repeated calls aggregate under one name. The ready trace repeats the step names and totals in bootstrapSteps. Nested steps overlap, so their durations should not be added together.
  • OPENCLAW_GATEWAY_RESTART_TRACE=1 logs restart trace: lines for restart signal handling, active-work drain, shutdown phases, next start, ready timing, and memory metrics. Ordinary stops also start a fresh trace with stop.signal.received and stop.drain timing. Named shutdown steps and coarse close phases emit .begin before waiting, then a duration when they settle; an unmatched begin identifies an entered phase that has not settled. These phases do not time every nested cleanup operation individually.
  • OPENCLAW_DIAGNOSTICS=timeline with OPENCLAW_DIAGNOSTICS_TIMELINE_PATH=<path> writes a best-effort JSONL startup diagnostics timeline for external QA harnesses (equivalent to config diagnostics.flags: ["timeline"]; the path is still env-only). Add OPENCLAW_DIAGNOSTICS_EVENT_LOOP=1 to include event-loop samples.
  • pnpm build then pnpm test:startup:gateway -- --runs 5 --warmup 1 benchmarks Gateway startup against the built CLI entry: first process output, /healthz, /readyz, startup trace timings, event-loop delay, and plugin lookup-table timing.
  • pnpm build then pnpm test:restart:gateway -- --case skipChannels --runs 1 --restarts 5 benchmarks in-process restart on macOS or Linux (not supported on Windows; restart requires SIGUSR2). Uses SIGUSR2, enables both traces in the child process, and records next /healthz, next /readyz, downtime, ready timing, CPU, RSS, and restart trace metrics.
  • /healthz is liveness; /readyz is usable readiness. Treat trace lines and benchmark output as owner-attribution signal, not a complete performance conclusion from one span or sample.

Without tracing, stops and restarts report nonzero active-work category counts at the first drain snapshot and at most once every 30 seconds while still pending. These reports omit task identities and request origins; categories can overlap. An ordinary stop logs active-work drain settled; beginning server close before teardown, including after a drain timeout or failure. Diagnostics do not change drain budgets or the service manager's stop deadline.

A client disconnect leaves interactive setup available for reconnect. A Gateway stop or restart closes setup prompts before draining work. Settings writes already in progress may finish, but setup will not wait for another answer during shutdown. After the Gateway starts again, reopen setup and check the saved settings.