Gateway and service recovery
Recover the Gateway service, a remote Gateway, Control UI assets, and Gateway tokens
These sections cover repairing the Gateway service, its remote and Control UI prerequisites, and the credentials it starts with.
Gateway service recovery
Run openclaw gateway status --deep to inspect the installed service and its
runtime before choosing a recovery action. Use openclaw gateway install for a
missing service, openclaw gateway start for an installed service that is not
loaded, or openclaw gateway install --force from the intended installation to
replace its service definition. Externally managed services still belong to
their supervisor.
Doctor also compares the service's package path and version with the active CLI,
without requiring a Gateway connection. Update finalization and standalone
openclaw doctor --fix reconcile eligible, previously running managed services
through the native installer; update-time Doctor reports drift and defers publication
to finalization. Doctor can automatically refresh installation-only drift in a
verified, writable packaged service. It also repairs recognized stale native
policy, such as a missing systemd KillMode=mixed or zero Scheduled Task restart
retries, through the update installer's backup transaction before restoring a
Gateway stopped for maintenance. Doctor reports changed keys and backup paths;
supported custom settings survive the rewrite. Automatic native-policy repair
preserves unknown operator edits and uncertain definitions for operator review.
Other command or credential changes still require interactive confirmation.
After successful standalone openclaw doctor --fix, an already stopped managed
Gateway starts and verifies readiness when its service targets the current
installation and final inspection positively verifies its ownership and offline
state. If that inspection fails, times out, or leaves ownership uncertain, Doctor
records the reason and leaves the service stopped. Inspect it with
openclaw gateway status --deep before starting it manually.
Update-time Doctor leaves activation with the updater. A stopped
service targeting another installation keeps its definition and stop state; run
the reported profile-aware openclaw gateway install --force command from the
intended installation to reconcile it (installation may start the service).
It preserves the service's profile and an explicit service port when no port is
configured. Source checkouts, deployment-owned overrides, and unavailable native
inspection do not grant automatic installation repair authority; Doctor reports
the mismatch and the next repair action.
If Doctor loses maintenance ownership during installation, it stops further installation or activation and restores its captured service definition when it can verify ownership of the replacement. The warning reports whether the definition was unchanged, restored, or needs inspection; follow the reported status and installer commands after the active maintenance or update finishes. Unverified restoration keeps recovery pending instead of claiming a safe restart.
When explicit repair stops a managed Gateway, Doctor waits for that process to
release shared-state lifecycle ownership within the service stop deadline before
repairing state. If ownership remains held, Doctor warns, restores the service,
and refuses the unsafe repair. On macOS, failed activation attempts restore the
LaunchAgent registration so its KeepAlive policy can recover; the error reports
whether the job is loaded and gives a recovery command if bootstrap also fails.
An ambiguous kickstart error followed by a check that confirms the job is absent
uses bootstrap recovery; successful activation then completes normally. A failure
for a job that remains loaded stays visible.
Doctor rechecks update admission after acquiring both maintenance coordinators. If it must cancel before repair starts, it reverses its own stop while its native service custody remains valid. Normal post-repair restoration still requires current update admission.
During service inspection, Doctor reuses already recorded update continuation receipts after rechecking live update ownership. This avoids repeating full database integrity scans for unchanged receipts within the service manager's inspection deadline. New continuation or takeover receipts still use the validated ledger writer.
On Linux, nested service-manager inspections share deadline accounting that excludes synchronous ownership and admission checks. Slow database inspection therefore does not consume the manager I/O allowance. Every authority check still runs; manager calls, cancellation, and service stop and drain deadlines remain bounded. Database inspection can still increase the total command time.
After restoration, Doctor verifies the Gateway with the shared health probe used
by openclaw gateway start. A running process or owned listener alone does not
mean it is ready. If the readiness budget expires while the service is running,
Doctor leaves it running and warns that readiness was not verified, including
the observed startup phase when available. Run openclaw gateway status --deep
or openclaw gateway diagnostics export to investigate; repair still completes.
When an exited container leaves a Gateway lock in a bind-mounted state directory, Doctor uses Gateway startup's namespace and heartbeat policy. It waits up to 95 seconds for an unverifiable container owner, bounded by any active service stop deadline, and reclaims the lock once its last renewal is more than 90 seconds old. An already stale lock needs no wait. The waiting message names the holder's host, PID namespace, and renewal age. If the holder keeps renewing, Doctor leaves state unchanged and reports how long to wait before retrying. A verified live holder on the same host must stop before ordinary repair can proceed; Doctor names its PID. Update-driven foreground shutdown retains its existing drain budget. Do not delete a lock that an active Gateway might still own.
Once Doctor owns maintenance, it checks the update requester's current authority within that same ownership scope. Its own maintenance window does not trigger a lock wait or an offline-maintenance refusal; requester revocation still stops repair.
When maintenance cannot acquire state ownership, Doctor includes the underlying schema or filesystem error. A shared-state database from a newer OpenClaw build stays unchanged; rerun Doctor with a build that supports that database version.
If Doctor's output pipe closes (for example, openclaw doctor --fix | head -20),
or Doctor receives SIGINT, SIGTERM, or SIGPIPE during maintenance, it waits for
admitted repair work and service restoration before exiting. An ordinary repair
error also restores the managed service Doctor stopped, using the current saved
configuration. Pending approval prompts cancel without interrupting admitted
writes. Concrete data risks, lost service authority, and unverified child
cleanup still prevent unsafe activation and report the recovery action.
For legacy services or conflicting systemd scopes, run openclaw doctor
interactively to review the findings and confirm supported cleanup. Cleanup
reports what it removed or skipped; it does not guarantee a replacement service
will be installed. Explicit repair maintenance skips this separate cleanup flow.
If Doctor stopped a managed Gateway for repair, a failed or timed-out restoration check produces a warning and Doctor still attempts to start that service and verify readiness. Live maintenance custody and update admission still apply; observed changes to the service command, account, or manager require operator review. An explicit ownership refusal is reported as a refusal, without attempting to start the rejected service. On systemd, Doctor retains the native manager and unit identity before stopping the service and revalidates it at activation. If that identity cannot be captured, Doctor leaves the service running and reports the inspection warning; live state writers still prevent unsafe offline repair.
When service inspection blocks repair, Doctor and gateway status --deep name
the failed native check:
- Linux inspection deadline expired: the manager check or its custody/admission
guards exhausted the inspection budget. This does not mean the user session bus
is missing. Check the reported restoration result and run
openclaw gateway status --deepafter recovery. - Linux user session bus unavailable: check
XDG_RUNTIME_DIRandDBUS_SESSION_BUS_ADDRESSfor the service account. A workingsystemctl --usercommand alone is insufficient: effective service inspection also usesbusctl --user. On Debian/Ubuntu, installdbus-user-session, then runsystemctl --user start dbus.socketfrom that account's user session. - Check cannot start (
EACCES/EPERM): check executable permissions and directory access as the service account. Native checks run from the filesystem root so an inaccessible operator directory inherited throughsudo -udoes not prevent inspection. - macOS GUI domain unavailable: sign in to the desktop as the target user
before managing its LaunchAgent. Error 125 for
gui/<uid>does not establish that a system LaunchDaemon exists. - macOS system domain unavailable or system LaunchDaemon detected: have root
inspect
sudo launchctl print system/<label>and stop the custom daemon through its deployment owner. To retain that supervisor, run Doctor as the state-owning account with the existingOPENCLAW_SERVICE_REPAIR_POLICY=externalpolicy. Keep the sameHOME,OPENCLAW_STATE_DIR, andOPENCLAW_CONFIG_PATHselectors used by the service. See Existing system LaunchDaemons.
OpenClaw does not manage custom system LaunchDaemons. Running Doctor as root
with another account's HOME does not add that capability and can create
root-owned state files.
For either platform, when an external supervisor owns the Gateway, have that
owner stop it and run Doctor as the state-owning account with
OPENCLAW_SERVICE_REPAIR_POLICY=external. This existing policy skips native
maintenance inspection and service mutations; it retains Gateway/state
coordinators and agent-database lease checks. Shutdown and restart remain with
the deployment owner. A failed native check is never treated as proof that the
Gateway is stopped.
Health diagnostics also leave native service inspection to that external owner. They still check the selected port, live Gateway ownership, and startup migration activity. With none present, Doctor reports the unavailable Gateway promptly instead of waiting for an unrelated native service manager. A live or starting Gateway retains the shared readiness budget.
For a system template such as [email protected] with User=%i, inspection
follows the current account's instance (openclaw@<user>.service) while
preserving the shared template. Run Doctor as that account after the system
service owner stops its instance.
Doctor waits for a starting local Gateway using the shared 60-second readiness budget, both on its initial check and after an approved restart. It reports the observed startup phase while waiting. A Gateway that still reports startup at the deadline produces a non-failing “still starting” result; Doctor leaves it running and does not offer another restart. Connection failure without startup evidence remains a diagnostic failure. This also applies when an installed updater invokes the candidate Doctor.
Plugin initialization and database startup checks can make a cold start take longer than ten seconds on a loaded or older host. Let the existing Gateway finish starting before requesting a separate restart. Remote Gateway diagnostics keep using the configured remote target.
Remote Gateway recovery
With gateway.mode: "remote", a failed Gateway health check does not trigger
local service install, start, restart, or bootstrap prompts. Check the remote
URL, credentials, and SSH tunnel or network connection. If the Gateway itself
needs recovery, run service commands on the host that runs it. A loopback remote
URL can be an SSH tunnel; it does not make the Gateway a local service.
See Remote access for connection checks. Other Doctor config and state checks still follow the selected posture.
Control UI assets
For source installs, Doctor can build missing Control UI assets or rebuild stale
assets after protocol changes. Its manual build command includes the detected
checkout path (pnpm --dir <checkout> ui:build), so you can run the displayed
command from another directory. Use the complete command, including its quoted
path, rather than running pnpm ui:build in an unrelated project.
Packaged installs without UI sources receive reinstall guidance instead of a source-build command. Doctor does not download a source checkout to repair a packaged installation.
Doctor reports obsolete cache/control-ui-assets directories left by older
installs. openclaw doctor --fix removes them; the update command's Doctor pass
also runs this repair. Cleanup failures are warnings and do not abort Doctor.
Invalid Gateway tokens
Doctor flags active Gateway tokens that are blank or contain the literal string
undefined or null. The Gateway rejects these values at startup. To replace an
inline token, run openclaw doctor --fix --generate-gateway-token, then restart
the Gateway. For a SecretRef, rotate the external secret source instead; doctor
preserves its reference and leaves password, none, and trusted-proxy auth modes
unchanged. An absent token still uses the normal startup token generation flow.
Known redaction placeholders, including __OPENCLAW_REDACTED__, are also invalid
credentials. Doctor and gateway status --deep name the affected reference even
if an older Gateway process still works with its previous in-memory token.
For a store-backed Gateway token, run openclaw doctor --fix (or
openclaw doctor --generate-gateway-token). Doctor verifies a database backup,
regenerates the referenced value, preserves the SecretRef and the entry's current
secret/env kind and allowed hosts, and prints the backup path. A credential
changed during backup is preserved. Restart the Gateway,
then reconnect or re-pair devices with the new token.
Explicit token generation reports when it skips a healthy SecretRef. Other placeholder secrets require a real replacement from their provider; Doctor reports them without deleting their stored values.
In trusted-proxy mode, a redacted inline or environment-supplied optional password does not block proxy authentication. Startup, Doctor, and status warn that local password fallback is unavailable. Replace or remove the optional password and restart; Doctor preserves trusted-proxy mode instead of generating a token.
macOS: launchctl env overrides
If you previously ran launchctl setenv OPENCLAW_GATEWAY_TOKEN ... (or ...PASSWORD), that value supplies fallback credentials when local configuration does not supply one. A configured inline credential or active SecretRef takes precedence over its matching environment fallback. A stale fallback can cause persistent "unauthorized" errors when it is selected.
launchctl getenv OPENCLAW_GATEWAY_TOKEN
launchctl getenv OPENCLAW_GATEWAY_PASSWORD
launchctl unsetenv OPENCLAW_GATEWAY_TOKEN
launchctl unsetenv OPENCLAW_GATEWAY_PASSWORD