Cloud worker troubleshooting
Symptoms and fixes for profile advertisement, authorization, bootstrap, enrollment, and teardown
Symptoms you may see when dispatching to or running on a cloud worker, and the check that resolves each one.
Troubleshooting
-
Worker transcript commit failed; check Gateway logs.— the Gateway could not complete a transcript commit. Deterministic failures, including malformed database admission receipts, end the turn; inspect theworker transcript commit failedwarning for the underlying cause before retrying. Temporary admission failures, including a concurrent schema publication invalidating a receipt, reconnect and replay with fresh admission. A failure does not prove that nothing was persisted. Recovery from a dropped acknowledgment still replays the original commit identity to avoid duplicate messages. -
Crabbox profile setup failed,Crabbox node runtime preparation failed, orCrabbox node enrollment setup failedwithcoordinator read retry N/M reason=timeoutandcontext deadline exceeded— OpenClaw retries the same fixed lease command up to three times with backoff inside the existing phase budget when every output line is a coordinator timeout diagnostic. Any script output prevents retry. Exhausted errors retain the coordinator diagnostic and attempt count; inspect coordinator availability before dispatching again. -
provider=<backend> does not support fixed idempotent lease IDs— OpenClaw cloud workers need a Crabbox backend with fixed lease ID support. Select a compatible backend; do not remove--lease-id. This exact exit-2 refusal occurs before Crabbox requests a machine, so a fresh dispatch fails permanently without a cleanup request. Other exit-2 failures and refusals during replay retain possible allocation responsibility. -
A worker remains
destroyingwithhas no local claim,does not match a valid ... claim,no exact resource-bound local claim, orstrict claim identifier mismatch— a missing or mismatched local ownership record does not prove remote absence. Check the exact lease in the provider inventory and restore the correct Crabbox state or follow the provider's ownership-recovery procedure before retrying Stop. The wording “if an earlier stop verified absence, nothing remains to do” is conditional; OpenClaw cannot use it alone to clear an existing cleanup obligation. -
No cloud profile is advertised — run the
operator.read-scopedopenclaw gateway call environments.list --params '{}'. If the response has noprofiles, ask an administrator to validatecloudWorkers.profilesand inspect the provider plugin; configuration changes reload without a restart; withgateway.reload.mode: "off", Gateway config writes such as Control UI saves restart the Gateway, and direct file edits wait for a manualopenclaw gateway restart. This is a configuration or provider-activation problem, not an authorization result. -
Cloud destinations are hidden or an RPC is denied — cloud profile dispatch and profile-target moves require
operator.admin.operator.writecan dispatch or move to an eligible paired device, move to the Gateway, and reclaim a placement;operator.readalone can discover profiles but cannot start, stop, or move a session. Profile configuration, infrastructure pairing, Connect machine, raw environment lifecycle, directexecNodeexecution, incognito sessions, and arbitrary host or node paths remainoperator.admin. -
The selected runtime lacks cloud placement support — choose a model whose advertised runtime supports cloud placement. The bundled OpenClaw and Codex runtimes are supported; undeclared runtimes remain local-only.
-
Codex cannot use a cloud profile — verify that the profile advertises
remote-exec, the Gateway enables a trusted Codex plugin installation, andgateway.nodes.commands.allowincludescodex.exec-server.stdio.v1without a matching deny rule. Bootstrap supplies the cloud-node plugin automatically. Approve the exact node invocation when prompted. Codex does not require an available OpenClaw worker slot; a missing plugin or denied command must be corrected rather than bypassed with Gateway or SSH execution. -
The portal tool is unavailable on a worker — confirm the session uses OpenClaw
worker-turnon an enrolled node that advertises portal-stream support. Update older node bundles when necessary. SSH-backedremote-execplacements, including Codex sessions, do not run the OpenClaw worker tool loop; move the session back to the Gateway withsessions.movewhen a Gateway-hosted portal is needed. -
"Worker bootstrap requires Node.js on the leased host" — add a Node install to
settings.setup(see The setup command). If setup still reports an unsupported Node version after APT installs Node.js 24, an nvm-managed Node earlier onPATHmay shadow it. The Debian/Ubuntu setup recipe upgrades that nvm installation in place and reports the resolved Node path and version if the final check still fails. -
gh: command not foundon a cloud worker — install GitHub CLI insettings.setup(see the Debian/Ubuntu example in Configuration), or install it on the paired worker host. Crabbox developer images include it; the sealed worker bundle does not. -
Repository preparation reports
clone-failedorcheckout-failed— the placement error includes the Git stage, exit or termination status, and bounded, redacted stderr. Use that detail to check Git availability on the node, repository access, or network connectivity before retrying dispatch. Native Windows repository preparation enables Git's long-path handling because nested session directories can otherwise exceed the limit when partial clones create.promisorfiles. -
Prepared project Git verification fails — the error identifies the Git subcommand and whether it timed out, exceeded its output buffer, failed to start, exited unsuccessfully, or received a signal. Prepared-workspace Git operations allow up to ten minutes, matching seed verification, within the provider's overall command deadline. A
git fscktimeout can indicate slow reads from a newly started snapshot; an unsuccessful exit needs investigation of the worker's Git object store. Full integrity checks still run before cache reuse. -
Windows workspace transfer reports
Filename too long— update the Gateway and reprovision the worker. New transferred workspaces enable Git's long-path handling in their private repository so pack import, checkout, and later workspace capture can use deeply nested session paths. Global Git configuration and Windows registry settings are unchanged. -
AWS instance-role attestation fails — clear
aws.instanceProfile(andCRABBOX_AWS_INSTANCE_PROFILE, if set). The plugin installs a supported Crabbox automatically before AWS admission; useopenclaw doctor --fixto diagnose a failed managed installation. -
Dispatch or workspace recovery fails — inspect
environments.listandsessions.describe. A failed environment exposes its bounded environment error. A failed placement exposesrecoveryErrorplus its durable per-sessionterminalReason; the selected Control UI chat shows that terminal reason above the composer. When deeper diagnosis is necessary, an operator on the Gateway host can inspect the durable worker state read-only. Do not edit the state database to bypass lifecycle fencing. -
Node worker launch is rejected — the error names the failed launch, status, or cancellation command and includes the node's bounded, redacted diagnostic. When the node rejects a new launch request as invalid (
INVALID_REQUEST), the turn fails immediately without cancellation because the node never registered that launch; an invalid launch descriptor usually means the Gateway and node run different releases, so update the older one. If cancellation could not be confirmed, the error also includes the cancellation failure; inspect the worker's state before retrying. A failed placement preserves its recorded cause; build-update guidance appears only when that cause identifies a build mismatch. -
Archive fails — the Control UI hides the session immediately and restores it only when the archive itself is rejected. Active session work must finish stopping; repeated requests cannot replace an operation already stopping that work. An already-failed worker can be archived while provider cleanup remains pending, with its worktree and recovery records preserved. Worktree cleanup errors after the archive is saved are logged for later cleanup. A client timeout does not cancel an accepted request or prove its outcome; check the session's current archive state before retrying.
-
Crabbox setup cannot reach the lease — check the selected backend's networking and setup-transport requirements in the Crabbox provider reference. Correct Crabbox's configuration and rerun
crabbox doctor --provider <backend> --jsonbefore retrying. -
Crabbox command execution fails — the error retains the bounded, redacted runner diagnostic. An execution failure can occur after the process starts, including during output capture or cleanup. Bootstrap version verification and plugin activation errors include the exit code, signal, and sanitized stderr tail; use that cause to diagnose missing dependencies or a terminated process.
-
A Linux profile allocates another operating system — update the Gateway. Linux allocations explicitly request Linux, overriding an ambient Crabbox target or config default. Windows mode and macOS market selection still follow the resolved profile.
-
Crabbox inspection reports coordinator read retries — Crabbox owns read retries within a one-minute budget. OpenClaw allows two minutes per lease inspection, including one minute for process startup and exit on a loaded host; Machine0 retains five minutes for readiness. Provisioning reads also respect the remaining overall deadline. If inspection still fails, check coordinator availability and the lease with
crabbox inspect --provider <backend> --id <lease> --jsonbefore retrying dispatch. -
Crabbox worker stays in
destroyingafter its lease was never admitted — update the Gateway. On the next reconciliation, a normally exited stop reporting404/not_foundfor both the coordinator lease read and release completes teardown and releases warm-image allocation ownership. Both responses must name that lease. Recognized direct-provider exit-4 absence also completes teardown; authentication errors, timeouts, incomplete output, and failed releases remain errors. -
Session shows a reclaimed or suspended badge after being idle — this is expected when its profile sets
suspendAfter. The next message provisions a replacement worker, warm when an image exists. -
A warm image is unavailable — a new allocation can select cold provisioning before its choice is recorded. An already admitted allocation keeps its original cold/checkpoint choice through retries. If its checkpoint cannot be forked, resolve the provider error or stop that allocation before starting a replacement; retry does not switch images silently.
-
Warm-image migration or capacity blocks dispatch — run
openclaw doctor --fixfor legacy state and follow its exact cleanup guidance. For capacity, stop outstanding workers or resolve pending image cleanup withopenclaw crabbox warm-images; allocation choices and cleanup obligations are never evicted to make room. -
checkpoint mode must be auto, native, or archiveand a paused capture — the selected Crabbox provider, target, or coordinator does not support the requested native capture, and an older CLI did not report a definitive unsupported-capture receipt. Update Crabbox to 0.69.0 or newer, which includes Crabbox #2613. OpenClaw then records the latest refusal and continues provisioning; later workers use an existing compatible snapshot when one is available and otherwise provision cold (shown as Cold only in Snapshots). Capture attempts for that image key resume afterwarmImages.refreshAfterhas elapsed since the refusal. Another refusal refreshes the marker timestamp; a successful capture clears it. Once no images, allocations, or operations remain, maintenance drops the marker-only row from local state afterwarmImages.refreshAfterwithout a new refusal. A capture already paused by an older CLI still needs paused-capture recovery: that refusal happened before any checkpoint was created, but confirm the source lease and capture time in the Crabbox checkpoint catalog before using--recover. Setsettings.warmImage: falseon the profile as a workaround to stop capture attempts. -
Project image capture fails — the session reports the underlying provider diagnostic, such as a checkpoint quota rejection, before its capture recovery instructions. Resolve that cause before retrying. An unresolved capture still blocks enrollment until provider artifacts and the recorded capture have been reconciled.
-
Cloud bootstrap requests a rebuild — run
pnpm buildin the Gateway source checkout, then restart the Gateway and retry. The running build, its package metadata, and the built plugin outputs must agree; editing source or matching the displayed version alone is insufficient. -
Cloud bootstrap download fails — the error identifies the connection, TLS, HTTP-response, or body-transfer phase. Crabbox retries transient transport failures and HTTP 502/503/504 responses with short backoff, stopping after three consecutive failures that make no byte progress. An attempt that grows the retained partial file resets that failure count. Total work stays bounded by the existing core-sized setup command deadline, and each stalled connection or body transfer times out after two idle minutes. Interrupted downloads retain their partial file and request the remaining bytes with
Range: bytes=N-, so a slow or reset-prone proxy does not force every attempt back to byte zero. A full HTTP 200 response replaces the partial file; an inconsistent range discards it and retries from zero within the same no-progress limit. The complete archive must still match its declared size and SHA-256 before installation; integrity failures discard partial bytes and stop. Each bootstrap artifact token permits up to 256 serial serves, counting both completed and interrupted full or ranged responses. This fixed cap accommodates frequent resumes while bounding artifact replays, including when a buffering proxy resets its downstream connection. Concurrent requests receive HTTP 503 withtransfer_in_progress; these busy responses wait with backoff within the setup command's deadline without spending the no-progress failure budget. Other HTTP 503 responses still count toward that limit. Bootstrap and node-worker-bundle transfer tokens share the size-derived bootstrap operation window of 45–95 minutes and are revoked as soon as enrollment or the operation ends; node-worker-bundle transfer tokens permit only one serve. Ranges reduce bytes per serve; retries never extend these lifetimes, raise serve budgets, or bypass owner revocation. Each retry logs the phase, error code, attempt count, and consecutive no-progress failures. Integrity, identity, unsafe-path, TLS-pin, and HTTP 401/403/404/409/410 failures stop immediately. Adownload TLSreset happened before an HTTP response; check the worker provider's outbound policy and the Gateway's TLS endpoint from the worker, not only from the Gateway host. For an HTTP status, check proxy routing and download authorization. Adownload bodyerror means response headers arrived; inspect the interrupted transfer, local disk, or archive-integrity error. Use a Gateway origin permitted by the provider's policy; do not disable certificate validation or bypass that policy. -
Bootstrap download fails with
ECONNRESET/ connection reset from the worker — repeated connection-level failures with no retained bytes mean the worker could not reach the Gateway public origin named in the error. Probe that origin from inside a provider sandbox, for examplecurl -v --connect-timeout 15 --max-time 30 https://gateway.example.com/(replace the example with the reported origin; no artifact URL or token is needed). An HTTP response shows that the TLS connection succeeded; a reset before any response points to the provider's outbound network policy or the origin's TLS endpoint. Downloads from other hosts succeeding does not prove this origin is reachable. Some providers' egress filters reset TLS to quick-tunnel hostnames such as*.trycloudflare.com; this was observed on Daytona on 2026-10-03. Configuregateway.publicOriginon a hostname reachable from that provider, then retry dispatch. -
Node enrollment times out — the command now uses the core-sized bootstrap window for its downloads and installation instead of a fixed 15-minute cutoff. Larger artifacts and slow Gateway uplinks receive more time, including resumed transfers. Project preparation counts both archives because concurrent downloads share the uplink. The outer provision budget reserves this work before grants exist and retains separate connection-wait, diagnostics, and cleanup allowances. Pairing lasts for the live enrollment and its size-derived window, including the node connection wait, rather than expiring after ten minutes. Closing or timing out an uncompleted enrollment revokes its pairing credential and prevents it from pairing; replay issues a fresh credential for the same setup identity. Inspect the bootstrap download or install error, node process state, and bounded node-log tail included in the enrollment error. Verify that profile setup installed a supported Node.js release and npm, that npm can reach the dependency registry, and that the box can reach the Gateway's advertised TLS URL. Forward
/__openclaw__/worker-bootstrap/artifacts/<sha256>as well as the public worker/node WebSocket routes through your proxy. If the error containsproxy_attribution_required, add the reverse proxy's source address togateway.trustedProxies. -
Client timeout while dispatching —
openclaw gateway calldefaults to a 10s timeout; pass--timeoutgenerously. Dispatch keeps running server-side either way, and an identical retry on the same Gateway joins that in-flight operation instead of provisioning another worker. A retry with a different profile or session identity is rejected. -
Provider authorization fails after
doctorpasses — read-only readiness does not prove permission to allocate or tear down a lease. Inspect the denied action and follow the selected provider's provisioning and cleanup requirements in the Crabbox provider reference. -
Worker runtime updating after a Gateway update — OpenClaw installs the current worker bundle on the existing machine, retaining its workspace, installed packages, and desktop. Background reconciliation starts installing the new runtime on reconnected paired hosts with retained sessions, normally within about a minute of reconnect. Follow that install in the Gateway logs, which record the start, progress every 30 seconds, interruptions with a reason, and completion; session RPCs such as
sessions.describeandsessions.listcan wait until the reconciliation pass finishes. For installs performed by a first dispatch or a submitted turn,workerRuntimeInstallin provisioning or active session placements and the Control UI chat notice also show transfer and installation progress. A turn submitted meanwhile waits for that same transfer; stuck-session diagnostics reportworker:runtime_refreshfor that waiting turn andworker:runtime_installfor a first dispatch that performs the install. Failed runtime updates retain the machine for recovery; inspect the Gateway's worker-environment logs for the installer error. Interrupted node turns resume through normal session recovery with fresh execution authority. Explicit Stop and Move requests still finish their teardown, and machines already lost by the provider cannot be reused. -
Cloud worker finished, but its workspace result could not be reconciled— read the detail afternode workspace command failed (<code>)and the matching warning in the node-host log. Workspace errors preserve bounded, redacted diagnostics, including filesystem and process-helper failures. If a root-owned container reportsworkspace quiescence refuses root-owned worker sessions, update the node host as well as the Gateway; the native Linux helper supports root-owned sessions without freezing other root processes. The older detached script still requires a non-root worker account. -
Cloud workspace conflict notice — the turn completed and kept the local version of each listed path. Use the staged-ref commands in the notice to inspect or take the cloud version; no retry is required for the non-conflicting changes, which are already applied.
-
Cloud session disk-space warning — delete unneeded files from the remote workspace or stop the cloud worker before large writes. The warning clears automatically after the next successful sample shows enough free space; a failed sample leaves the last successful warning visible and does not affect the session lifecycle.
-
“The previous cloud turn's workspace result is still reconciling” — the Gateway waited briefly for the prior result's durable fence and could not acquire the session claim. Wait for reconciliation to finish, then retry the turn; restarting the Gateway is safe because recovery preserves staged results before reclaiming a dead worker.
-
GitHub publication failed — for Gateway-brokered publication through Publish PR or remote-exec
github_publish, open Agents → Tools → GitHub Identity and confirm the effective@login, selected scope, access expiry, and refresh state. Reconnect GitHub when refresh is expired or unavailable; use a managed PAT only as the explicit fallback. For push rejection, inspect repository write access and branch drift;/userverification does not prove repository write access and the broker never rewrites published history. If the published branch no longer fits the accepted history, preserve local work and apply the intended edits on top of the published head to refresh the existing PR, or use a new session branch and replacement PR for intentionally rewritten history. Do not repeatedly publish the same divergence or merge old history merely to make a rebased branch pushable. Failed branch observation requires restoring read access/connectivity, not assuming the branch is absent. For pull request rejection, grant pull-request write access and retry Publish PR or callgithub_publishagain with a new tool call. -
Repository publication is unavailable — Git clean filters, unsafe Git configuration, or failed publication-snapshot validation can prevent preparing a publishable checkpoint. Raw recovery checkpoints and normal Stop still preserve accepted changes. Correct the repository configuration, then run another turn or save an edit to prepare a new checkpoint before requesting publication again.
-
Lease housekeeping —
crabbox list --provider <backend> --jsonis a read-only inventory.crabbox stop --provider <backend> --id <lease>andcrabbox release --provider <backend> --id <lease>are destructive and release a lease manually. OpenClaw keeps the lease alive while its session is placed, then stops heartbeating during teardown so genuinely idle leases expire on the profile'sidleTimeout. A temporary heartbeat failure, including a lease claim conflict, warns and keeps the next scheduled renewal. Backends that do not support lease heartbeat produce a warning; the plugin ensures the CLI itself supports the command before use.