跳到正文
FunCoding

搜索

搜索文档、文章、Skill 和 MCP

Update repair and recovery

Recovering from a failed openclaw update, plus the update repair and update cleanup subcommands

What happens when an update fails, and the subcommands that finish the job. Part of the openclaw update reference.

Recover a failed update

After a failed interactive update or repair, OpenClaw finishes cleanup and offers Diagnose update failure, Report update failure, or Exit. Reporting previews the sanitized issue body and requires separate confirmation.

Unexpected exceptions retain the known update mode, resolved target, failed step, and any recorded recovery outcome. Reports include a bounded, redacted error code or name and first message line through the same diagnostics as failed commands; unrecognized private text and stack traces are excluded from the public preview.

Choosing Diagnose update failure opens Triage, which starts the first directly launchable coding agent on PATH, in this order: Claude Code, Codex, OpenCode, then Pi. It passes the captured update failure directly and leaves fresh Doctor checks and diagnostics collection to the agent, so a broken installation does not delay the handoff. The agent keeps its existing authentication, sandbox, and approval settings.

The agent starts in the operator's original working directory, or their OS home if that directory is no longer accessible. The failed installation's resolved state, config, and default workspace paths remain pinned for the repair.

A verified rollback leaves the previous generation running. The interactive menu selects Exit by default; declining or cancelling keeps the failed update's nonzero exit status. After a verified rollback, --json, --yes, non-interactive, and managed-service handoff runs do not prompt or collect automatic diagnostics.

For failures without a verified rollback, updates using --yes, --json, or a non-interactive session (including piped input or output) collect diagnostics and print handoff commands without starting an external coding agent. Eligible failures can start post-failure triage on configured inference after update ownership and service compensation settle. With --json, triage output goes to stderr so stdout retains the original update result. Diagnostic collection failures never hide the update failure.

For a background or Control UI failure, use the installation-specific command printed on the Gateway host. Printed commands use PowerShell on Windows and POSIX shells on macOS, Linux, and WSL. When running triage manually, keep the same profile and state/config overrides:

openclaw triage
openclaw triage --agent codex

Use openclaw triage --non-interactive to collect diagnostics without starting an agent. Add --update-result <path> to include a saved update-failure artifact.

When another process saves configuration during database admission, OpenClaw warns and reads the current configuration again. It validates and uses that configuration before continuing, retaining concurrent changes when applying the requested update. If the root config or an included file changes after candidate checks, it repeats those checks against the current configuration before activation. A candidate that cannot accept the current configuration still fails validation; a concurrent save alone is not a refusal. If the save changes an implicitly selected update channel, OpenClaw resolves the target again before execution. An explicit --channel keeps its selection.

Validation failures leave the serving Gateway untouched. If stopping the managed service unloads it and then fails before activation, OpenClaw attempts to restore the verified original runtime after rechecking service ownership. After activation, a failed verification can restore the previous package when database schemas are unchanged and the config file still matches the candidate’s activation Doctor output. Preserve migrated state and history; replacing the code alone cannot undo a migration. The original failed update still exits nonzero after the agent finishes, even if the repair succeeds.

On Linux, systemd can unload an inactive unit after the updater stops it. The owning updater reloads that unit's metadata when rechecking admission, retaining the original manager and service identity. This does not start the service or rewrite its definition. A later refusal still uses the recorded stop to restore the previous Gateway; a service that was already stopped remains stopped.

An already-current installation can still need plugin maintenance. If the update parks its Gateway for Doctor and a maintenance step is refused, recovery uses the current run's latest stop and restarts the installed package after Doctor's writers settle. It records the observed recovery outcome while preserving the failed update result. Earlier runs' stop receipts never authorize that restart; operator-stopped services and explicit data-risk refusals remain stopped.

After activation succeeds, a failure to read or publish update reporting leaves the updated installation in place. Reporting failures do not trigger package rollback. The command still exits nonzero when required finalization cannot complete; follow its recovery guidance after the owning updater exits.

Activation Doctor rechecks the chat requester's authority inside its own live maintenance scope. This lets it read authorization policy while the state database is offline for repair, without granting access to unrelated operations. The candidate supplies this repair even when an older updater launches Doctor.

Git updaters with database rollback support snapshot the stopped installation's databases before activation Doctor. If Doctor fails, verified snapshots restore the pre-migration state before the previous source, runtime, configuration, and managed Gateway are restored. Doctor records its admitted and settled database generations against that backup, so its own migration does not prevent rollback. Later operator or runtime writes prevent destructive restoration: the updater preserves them, restarts and verifies the Gateway on the migrated installation, and reports the refusal with the next Doctor command. Lost ownership still prevents unauthorized effects and retains the recovery artifacts. This rollback support belongs to the installed updater; a new candidate cannot add it to an older updater already running.

Dry runs and commands rejected by the initial argument, external-supervisor, state-store ownership, handoff identity, or immutable-config checks do not collect diagnostics or start an agent. Once those checks pass, failed metadata, schema, runtime, and managed-service checks use the failure actions above even when installation is blocked. This includes an update that cannot safely stop its parent Gateway process. Diagnosis preserves that refusal: it does not stop the Gateway, retry the update, or bypass safety checks. See Update troubleshooting.

Original state captures

Before a fresh direct CLI update writes runtime state, the installed updater attempts to retain the original config, its includes, local databases, and declared plugin migration resources. The capture stays beside the selected state directory in <state-directory>.update-captures/<run-id>. Doctor continuations keep the same capture; they do not replace it with already migrated state. State-directory relocation leaves its original location and recorded paths intact. Inherited control-plane and managed-helper runs retain their existing capture behavior. Standalone doctor --fix preserves a separate pre-repair copy; that copy does not replace an earlier update's originals. Standalone doctor --fix snapshots databases under its own maintenance custody with a single isolated backup worker. Shared database families larger than 64 MiB keep discovery copies in an isolated process so a slow copy does not block Doctor's main thread from processing cancellation. Standalone Doctor captures are retained for 30 days: the next standalone doctor --fix retires older sealed Doctor captures and reports each removal; incomplete captures and update captures are never retired automatically, so take a verified backup when you need a long-term copy. Review and retire superseded update captures with openclaw update cleanup.

On Linux filesystems that reject native no-replace rename, fs-safe uses exclusive hard-link publication followed by source removal in native auto mode. Existing captures are never overwritten; native require mode still refuses unsupported publication. If publication stops with both names present, OpenClaw retains the capture as incomplete evidence. A remaining manifest.json.partial prevents reuse or automatic retirement even when manifest.json contains complete JSON. Keep both names for manual inspection; their presence does not authorize restoration.

These captures are evidence for manual recovery. Active writers can change state during capture; an observed change leaves the capture incomplete and produces a warning. The set is not an atomic snapshot across active stores. Missing, unreadable, or incomplete captures do not establish a safe rollback point. The updater process keeps optional debug-proxy persistence disabled because its update history can use an older database schema. Doctor can resume capture after preserving the originals and admitting the repaired schema. A successful update that skips Doctor can therefore leave local HTTP tracing disabled for that invocation. Direct updates, including --dry-run, report this limitation when debug capture is enabled.

Use openclaw update status --json to inspect retained evidence. Runtime rollback does not prove that an earlier original capture was restored. Status reports that capture as restored only when the restoration evidence identifies its manifest. Standalone Doctor copies appear as manual; their presence does not record a successful or failed repair. An unfinished capture appears as incomplete alongside valid captures, with its directory and no sealed manifest reference. Keep current data and inspect the originals before attempting restoration. Older installed updaters may not preserve or forward an original capture; a newer Doctor reports that limitation instead of treating current bytes as the pre-update state. If capture discovery cannot verify an older driver's history, Doctor warns and continues repairs under its existing maintenance and update ownership. Required migration backups still apply. Take a verified backup before an upgrade when you need a complete recovery copy.

Inspect migration preservation

When both the original and post-migration state have sealed recovery captures, compare them without opening live databases:

openclaw database verify-preservation /path/to/original/manifest.json /path/to/candidate/manifest.json \
  --original-sha256 <recorded-original-manifest-sha256> \
  --candidate-sha256 <candidate-manifest-sha256> --json

Use the original digest recorded before migration. The command verifies each capture's manifest and payload hashes, then derives versioned semantic witnesses from the original inventory. A fresh capture cannot replace that original reference. Retained captures must remain at their recorded locations.

The first projection supports agent schemas 24 and 25, the current shared-state schema, and unchanged declared plugin SQLite stores. The 24-to-25 transition allows the migration owner's validation queue, canonical receipt, and schema metadata changes; session generations, retained transcript bytes, snapshots, and plugin rows must survive. Retained schema objects, row identities, SQLite application IDs, and file modes are included. Other transitions fail as unsupported. Shared registry version and observation fields may refresh only for those exact agent stores after their content and target schema have been verified; registration identity and unrelated rows remain exact. Recorded file symlinks are resolved from the sealed inventory. The existing capture format does not retain external parent-directory alias mappings; inspection refuses those bindings rather than guessing from today's filesystem. Protected files remain byte-exact. Unchanged missing resources and classified history gaps produce preserved-with-warnings; changed or unreadable evidence exits nonzero.

This is an explicit cold inspection, with at most 4096 resources and 512 tables per database. It streams all retained rows and may take time on large captures. It creates no new backup or witness store and does not authorize migration, rollback, startup, or activation. Updates retain their existing safety gates; automatic migration acceptance is a separate integration.

Retained updater runtime

An update can retain its running code in an openclaw-update-runtime-* directory beside the installation or in the system temporary directory. The updater settles its workers after success, failure, an exception, or SIGINT/SIGTERM, including failures while reporting the outcome. Complete projections record Runtime retained at <path>: <reason> for the next eligible update or Doctor cleanup, so recursive deletion does not delay command exit. Incomplete preparation keeps immediate best-effort cleanup. Unsettled workers retain their runtime with the failure reason. A cleanup warning does not replace the original update outcome. An earlier nonzero exit remains nonzero while workers are draining. Mutation and recovery owners must still drain; their failures produce a nonzero exit even if the printed command result was successful.

Retention copies plugin manifests and files inspected by plugin safety checks, so retaining the updater does not make the checkout's plugins fail hardlink validation. Other runtime files remain hardlinked when supported.

Updaters using @openclaw/fs-safe 0.23.1 or later rely on its guarded byte-copy fallback when a container or filesystem refuses file cloning. Native filesystem safeguards, source identity checks, file modes, snapshot verification, and SQLite byte-copy space admission remain active. Genuine I/O errors still fail the copy. This includes Proxmox LXC containers whose seccomp policy denies the FICLONE ioctl; changing that policy is unnecessary for an updater using this dependency. Do not globally disable native filesystem support to bypass cloning: Doctor's state migrations require native safeguards.

These lifecycle and copying changes apply when the installed updater supports them; installing a newer candidate cannot change the updater already running. For the first hop from 2026.9.7 in an affected container, manually install a release containing this fix with npm. Subsequent openclaw update runs inherit the fallback from the installed updater's fs-safe dependency.

On Windows, interruption before activation still lets the admitted recovery owner restore task autostart after pending task operations settle. Cancellation fences new update work; restoration still requires the original live installation owner and verified task ownership. After that updater exits, run the newer openclaw doctor --fix from the original checkout to locate its sibling runtime directories. Doctor also checks known temporary directories, including the managed service's TMPDIR, TMP, and TEMP. Relative service paths require its recorded absolute working directory. Legacy pnpm projections remain recognizable after a versioned package-root change within the same store; lookup is limited to 4,096 immediate projected-store entries. Recognized runtime projections are disposable; Doctor removes them when no worker still uses them. If ownership or process liveness cannot be verified, Doctor preserves the directory and reports the reason.

Candidate Doctor stack overflow

Chat-triggered updates to 2026.9.6 can fail with authority-check-failed: Maximum call stack size exceeded, sometimes preceded by Update history reconciliation could not complete. This is a candidate Doctor authority-check defect; it can happen on the first update, without migrated state or earlier failed runs. The corrective candidate can run through the installed updater with retained history intact. Running the older installation's standalone Doctor cannot fix code in the candidate package.

There is no supported command to reset retained update history. update repair finishes interrupted finalization, and update cleanup retires eligible recovery originals; neither clears the run ledger. Keep history and backups rather than deleting database rows to work around this failure.

Docker image-layer package updates

Docker OverlayFS can reject moving an npm package installed in an image layer with EXDEV, even when the package and its backup are on the same mount. Updaters with the copy fallback immediately retain and verify an independent copy, record its identity for recovery, and remove the original before publishing the candidate. A mismatched copy leaves the original package intact. The verified copy remains available for rollback until activation is confirmed.

The first update from 2026.9.7 or 2026.9.8 still needs a manual installation hop: those installed updaters run their own publication code, so the candidate cannot supply this fallback. Prefer rebuilding the Docker image with the desired OpenClaw version. For an in-container replacement, follow the manual update precautions, including a verified backup and stopping the managed Gateway, then run npm install -g openclaw@<version>. Run Doctor and, if needed, openclaw update repair from the new installation before restarting through the service manager.

update repair

An older updater can leave package activation at prepared after refusing an update before publication. Run openclaw update repair from an installation containing this fix. Repair verifies that the original package and launchers are unchanged, aborts the unused preparation, and retires its recovery artifacts so the next update can proceed. It does not publish the staged candidate. An active updater, changed package or launcher, or unfinished state restoration keeps its existing recovery checks. A candidate cannot patch the older updater already running; use the manual installation hop below if the installed CLI lacks this repair.

For a publication stranded at publishing after an external write, repair can close it as publication-settled-external-change when the installed build-info reports the exact candidate version, every file in the package's own dist content inventory still matches, and no updater owns the installation. Repair uses fresh executor ownership even if the old lease store was removed or replaced. It never executes the old helper; changed or missing helper bytes and replaced retained directories are not required proof of the live candidate. The root package.json must parse with name openclaw, the candidate version, and type module; every main, exports, and bin target must resolve to a file in the package. Targets within dist/ must be inventoried; top-level targets such as openclaw.mjs are checked for resolution without content verification. Extra package.json files under dist/ refuse settlement because they can change how inventoried code loads. Dependency manifests under node_modules/ are expected and ignored. Other extra dist files remain warnings. Restore any changed inventoried file to its packaged bytes before retrying; a working Gateway alone does not waive an inventory failure. Repair preserves the previous package and any remaining helper, leaves the installed package and launchers in place, and records the warning and extra paths in update history. The warning and receipt identify the root manifest as field-verified, not content-verified. The sealed tree digest cannot identify old per-file metadata differences. Use a CLI containing this fix; the original sealed helper keeps its original recovery checks.

After recording a settlement, repair moves the completed control journal intact into control/ inside the reported recovery-evidence directory. It no longer appears as active packageActivation state or blocks an older updater on an unfamiliar settlement reason. If reporting stops before that move, rerun repair from the compatible CLI. Archival failure after durable completion is a warning: evidence stays preserved and does not prevent repair finalization or a later update. Retained evidence is not deleted or used as authority for later updates. Unfinished operations still require a compatible recovery owner. This behavior does not deliver a newer repair implementation to an already-blocked older CLI. For completed anchor-retired history, use the independent helper recovery to preserve the receipt and unblock the original updater without replacing its installation. Unfinished recovery retains the first-hop installation limitation.

For a package update stranded by an older updater's launcher ownership checks, use the manual installation hop, then repair from the new CLI at the same root:

npm install -g openclaw@latest
openclaw doctor --fix
openclaw update repair

Follow the manual update precautions, including a verified backup and stopping the managed Gateway during replacement. When the installed package directory matches neither recorded generation, repair closes the previous package operation as superseded-by-manual-install, warns with its operation ID, and preserves any remaining staged files and helper beside the installation. Changed or missing historical artifacts and archive collisions become maintenance warnings after the durable close; they do not prevent finalization. The original failed history entry remains intact. The pending package-recovery gate then clears, so another update can proceed. Unfinished restoration keeps its original recovery checks; missing packages, active update owners, and pending database or configuration restoration still require their existing recovery path.

If recovery instead reports managed handoff lease database identity changed, run openclaw update repair from a CLI containing this fix. Repair acquires fresh update ownership on the current lease database. A still-installed published candidate must pass the verification above before its operation can close. An untouched obsolete preparation can close as recovery-lease-identity-changed. It warns with the old operation ID and retained artifact path, leaves the installed package and launchers in place, and clears package admission for the next update. The original helper cannot recover against a replaced lease database. Matching lease identities keep the original recovery checks; another live update owner still prevents settlement. No recovery artifacts are deleted. An older installed CLI cannot obtain this fix from a candidate it has not yet staged; use the manual installation hop above.

The same repair handles ENOENT when the recorded handoff lease database is missing, for example after a reboot clears a temporary filesystem. Its storage owner recreates the lease database, and repair acquires fresh update ownership before closing an untouched obsolete preparation as recovery-lease-missing. A missing or changed lease never permits discarding an unfinished rollback. The installed package, launchers, and retained recovery evidence keep the same protections. Repair then continues through Doctor and plugin convergence; plugin data/settings warnings clear only when their migration owners complete the required work. Remaining warnings name the next repair action.

Rerun update finalization after the core package already changed but later repair work did not finish cleanly. This is the supported recovery path when openclaw update installed the new core package but post-core plugin sync, managed npm plugin metadata, registry refresh, or doctor repair did not converge.

openclaw update repair
openclaw update repair --channel beta
openclaw update repair --json
openclaw update repair --accept-capabilities

An interrupted automatic triage can leave an uncertain installation handoff after its updater and helper exit. Explicit openclaw update repair can reclaim that handoff when both recorded PID/start identities are provably dead, a complete host census finds no remaining references to its run or retained paths, and at least 45 minutes have passed since the lease's last recorded activity. A recoverable larger recorded timeout extends that grace period. Gateway startup and borrowed update processes do not reclaim these leases.

Repair also checks legacy update child-lineage leases left after a manual installation hop, even when the installation-root lease is absent. Dead PIDs and PIDs reused with different start identities can be reclaimed. An unbound lineage reservation names the same executor and helper; an unrelated process group with that numeric ID does not keep it alive. Bound child process groups still require proven extinction. Repair preserves the original lease evidence while claiming the installation and records the reclaimed keys in update history. When several child leases remain, repair checks each child's original run and rollback evidence before claiming the installation; one child's history cannot authorize reclaiming another child's lease. Live or uninspectable owners remain protected. On Linux with restricted /proc visibility, retry from the original OS account with process-inspection permissions; permission errors never prove that an owner died.

An ordinary installation-root update lease is not legacy custody just because it has no mutation-protocol marker. Repair leaves that lease to its current owner; normal update admission can reclaim it once its owners and descendants settle.

The original run must be identifiable from its retained helper, update history, or generation-bound repair metadata, and readable in the selected state database. Repair needs that record to check rollback and recovery evidence. Use the same profile and state overrides as the failed update; repair does not substitute a lease owner ID for a missing run.

Refusals name the processes, inspection gap, or remaining grace period. Stop named work through its owning terminal or service, then retry; do not delete the lease database. Repair retains unreadable helper paths in its census instead of assuming their work has stopped.

On Windows, identifying foreign process owners can require Administrator privileges. If repair requests elevation, use the same Windows account and preserve the failed update's profile and state overrides. Unknown ownership remains unverified.

When no rollback step is recorded, repair uses ordinary current-installation finalization and records a handoff settlement in update history before releasing ownership. It preserves retained artifacts and does not invent an owner-death time or previous Gateway state. Unresolved state restoration keeps its existing recovery safeguards. A failed repair retains its original evidence, bound to the new lease generation, for a later explicit repair. Each repair attempt records its own run before starting finalization work, so a later repair also checks for descendants of interrupted repair attempts.

An unfinished package or configuration rollback keeps the handoff and its artifacts intact. Inspect openclaw update status --json and complete the recorded restoration before retrying repair. Verified completed rollback and settlement receipts permit repair, including installations without a running Gateway. Skipped rollback steps and diagnostic warnings alone do not prevent repair.

When update, post-core continuation, or repair runs under Bun, its OpenClaw maintenance children use that same Bun executable, including fresh Doctor, config validation, readiness, completion, and non-interactive failure diagnostics. Managed-service runtime selection stays with the service definition and its runtime pin.

If an older updater publishes the new core but then reports update-executor-settlement-failed with Parent executor is suspended for its candidate., wait for that updater to exit and run openclaw update repair --yes --json from the updated installation, preserving its profile and state/config overrides. This finishes Doctor and post-core convergence through a fresh owner. Check the repair result before restarting an already stopped Gateway through its service owner. Updating the candidate cannot change the older updater already in memory.

When a managed Gateway was already stopped before standalone repair, repair leaves it offline and warns that you must run openclaw gateway start to bring it online. If its service definition points to a different installation, repair instead reports the installation repair command. These maintenance warnings also appear in postUpdate.doctor.warnings; otherwise successful finalization reports status: "warning" and exits successfully.

FlagDescription
--channel <stable|extended-stable|beta|dev>Persist the core update channel before repair. For extended-stable, eligible official npm and trusted official ClawHub plugins that follow bare/default or latest intent target the exact installed core version. Extended-stable repair is rejected on Git checkouts without changing config.
--jsonPrint machine-readable finalization JSON.
--timeout <seconds>Override each repair phase deadline in seconds. Defaults vary by phase (see below).
--yesSkip confirmation prompts.
--accept-capabilitiesAccept each plugin's reviewed capability changes while repairing plugin state.
--no-restartAccepted for parity; repair does not request update activation. The repair parent restores a Gateway it stopped for maintenance.

Untouched, identityless 2026.9.2-era update admissions heal automatically after more than 24 hours. Gateway startup, openclaw update status, and openclaw status retain the row as an abandoned failure with reason legacy-driver-expired and an advisory to run openclaw update to retry. The Control UI refreshes the Gateway's recovery classification before refusing a suspended config write, so an expired orphan does not keep settings or provider sign-in blocked. Live updates and pending recovery remain protected. No explicit repair is needed for this shape.

update repair first inspects stale update history. When the installed Gateway generation is healthy and the only remaining problem is an inactive ledger row, repair records failed / abandoned and exits successfully without Doctor, maintenance, or a service stop. openclaw status and the Control UI then report the abandoned run as reconciled, without a failure warning or retry prompt; its historical failure record remains intact. It also acknowledges a Gateway-reconciled row once within 30 minutes of reconciliation. Later repair invocations use full finalization, so historical recovery cannot suppress plugin convergence. Successful repair acknowledges every outstanding abandoned run in Doctor's latest 100 history records, regardless of age or whether a newer update exists. The original failed outcome remains recorded, including when its target build is unknown; Doctor stops repeating the repair instruction after acknowledgement. Explicit recovery does not wait 30 minutes when every recorded updater process is provably dead (its PID is gone or its process-start identity has changed). Identityless rows and runs with an unrecorded adopter still require more than 30 minutes of inactivity; a recorded live or uninspectable driver blocks abandonment recovery. JSON output identifies reconciled run IDs in reconciledRuns, with status: "ok", mode: "repair", and restart: false.

Repair also acknowledges an untouched package installation whose update was refused because its package-manager owner was unknown, once the installed version meets or exceeds the resolved target. This includes older updaters that incorrectly recorded that refusal as a failed update. The original refusal detail stays in history; repair clears the failure prompt without Doctor maintenance or a service restart. Runs that reached installation or finalization still require normal repair.

Fresh Doctor children run with the existing external service-repair policy because the updater owns service changes. They preserve an operator's OPENCLAW_SERVICE_REPAIR_POLICY=external selection and retain Gateway/state coordinators and agent-database lease checks. An external deployment owner still owns stopping and restarting its Gateway.

Automatic repair finishes its embedded agent turn and releases that turn's database and process resources before asking the update owner to run Doctor or update repair. This prevents the repair agent's own credential writes from blocking maintenance. Other live agent leases still block repair. Maintenance preserves the original Gateway activation intent, including --no-restart and intentional stops. A successful maintenance command alone does not verify the original symptom.

Repair invoked within the owning update can continue when its inherited run ID and live process identity match that owner. Standalone repair records the same continuation for its new run and passes that run ID to its Doctor children. The repair parent uses Doctor's maintenance lifecycle to stop the owned Gateway, then releases its database locks before running the Doctor children. The children repair state without stopping or restarting the service. The parent restores and verifies the same service after convergence, including when a Doctor child fails. An already stopped service stays stopped. Service ownership and the invoking run are revalidated before every native operation. A restoration failure names the cause and the commands to inspect and restart the Gateway. Normal update finalization continues to leave activation with its outer updater.

If Doctor reports that the update parent must stop the managed Gateway, wait for that update to exit, then run openclaw gateway stop and retry openclaw update repair from an independent shell with the same profile and state/config overrides. On macOS, stop unloads the LaunchAgent and verifies that its process exited. A still-loaded service or surviving PID after a successful stop is a service shutdown failure. If stop cannot unload the service, use the exact launchctl bootout command printed in the refusal from the owning user's logged-in macOS GUI session.

An unrelated update whose driver is live or cannot be inspected still blocks repair, even after a long period without activity. Manual doctor --fix also refuses to stop a service while that update is active. The refusal identifies the owning run, phase, driver PID, host, start and last-activity times and ages, and observed liveness (alive or not observed). Wait for that update to finish, or stop the named driver on its host and rerun openclaw update repair after it exits. Elapsed inactivity alone does not authorize taking over a live updater.

Explicit channel or capability changes and known incomplete post-core work use full finalization. Recorded activation, restart, verification, or finalization steps require that convergence even if the Gateway has already reconciled the run. Repair checks newer failed post-core history as well as active rows, including older finalization attempts with an empty failure reason; an older stale row cannot hide unfinished work from a newer update. If the bounded history inspection is incomplete, repair also uses full finalization. The parent parks its owned service before Doctor enters maintenance; a Doctor child cannot take service activation from an update parent. Successful full finalization then reconciles the selected stale rows and acknowledges unacknowledged abandoned outcomes from the 100 most recent history rows captured at repair admission. Those outcomes need not be the latest run or less than 30 minutes old; the time limit applies only to skipping full finalization. Additional historical outcomes require full finalization, even when the latest row qualifies for the lightweight repair. Historical failures and their details remain intact; acknowledgment clears their Doctor repair prompts, not their failed status. Rows outside the captured history window and new runs admitted during repair are not acknowledged by that invocation. Failed convergence leaves the selected rows intact. If any selected run resumes before reconciliation, the whole selection is preserved. Full finalization JSON includes reconciledRuns when rows were selected for recovery, listing the IDs newly acknowledged by that invocation, including already-terminal abandoned rows. Successful convergence with nonfatal warnings also acknowledges those rows. Deferred maintenance preserves the selected history and pending migration obligations.

For full finalization, update repair runs openclaw doctor --fix, reloads the repaired config and install records, syncs tracked plugins for the active update channel, updates managed npm plugin installs, repairs missing configured plugin payloads, refreshes the plugin registry, and writes converged install-record metadata. If plugin migrations remain deferred, finalization runs another fresh Doctor after releasing install-record ownership, even when no plugin package changed. This lets a corrected local plugin finish its pending confirmation in the same repair run. Doctor preserves the plugin's configuration when compatibility checks prevent discovery, so correcting the plugin does not require recreating its allowlist or enabled entry. Configured runtime plugins whose versions follow OpenClaw are checked against the newly installed core during post-update repair, even when the updater process started on the previous version. It does not install a new core package or request update activation. The repair parent restores a service it stopped for maintenance, as described above. Human output ends with a finalization result that distinguishes completion, completion with warnings, and failure.

When repair finds a configured npm plugin payload but cannot recover its install record, it reinstalls from the selected registry source, using the active channel or exact version pin. This requires registry access; if verification fails, repair preserves the existing payload and does not publish a new install record. Registry verification and any required capability review finish before the repaired install record is published.

When a bundled plugin moves to an external package, failed relocation reports that the replacement payload was not installed and preserves the underlying error. Resolve that error before retrying with openclaw update repair. Doctor and update repair reinstall configured payloads with missing package files or a reported missing runtime entry; an empty directory is not a successful installation. Rollback removes empty managed npm projects after staged files are cleaned up. Doctor preserves external companion packages and their install records even when a source checkout also contains a bundled-discovery copy of the same plugin. Repair diagnostics must identify the recorded package root; a broken same-ID source copy does not trigger replacement of a healthy managed package.

With --json, stdout contains one JSON document. Doctor panels and other diagnostics go to stderr, so stdout can be parsed directly. Plugin-only availability, installation, or load failures appear in postUpdate.plugins.warnings; finalization reports status: "warning" and exits successfully when required checks pass. Doctor maintenance admission refusals also finish with a warning when no data is at risk. Repair restores any service it stopped and leaves migrations pending. When Doctor could not run, the saved run is skipped with reason doctor-maintenance-pending, and the report leads with the next action: stop the Gateway through its service owner, then rerun openclaw update repair with the same profile and state overrides. A standalone repair's generic failure is recorded as repair-failed; specific failure codes and failures belonging to an existing update keep their original reasons. Errors after repair writes begin, a live or unverified Gateway, unreadable state, active migration writes, unsettled cleanup, invalid configuration, and failed required readiness checks still exit nonzero.

If a required repair phase exceeds its deadline, repair exits with code 1 and JSON reports status: "failed" with the stuckPhase. That result remains available after service restoration, including when the Gateway is still starting or restoration also fails. A startup warning after otherwise successful Doctor repair does not clear a repair phase timeout.

Recorded pending-migration warnings stop appearing after the migration owner records completion. Unrelated warnings and later or reintroduced obligations remain visible; the original update history is preserved.

After post-update or finalization work fails and its child processes settle, OpenClaw checks the installed Gateway using the normal startup and readiness budget. If maintenance found no Gateway service or listener, recovery records that readiness observation was skipped instead of waiting for a Gateway to appear. Package and database restoration checks still apply, and the original failure remains recorded. Update history and failure reports record the observed serving version and readiness. A standalone repair failure before Doctor maintenance begins uses one bounded observation because that repair has not requested Gateway startup. Observations also cover foreground Gateways. A failed finalization step can therefore report verified serving while retaining its original failure and repair guidance. The observation does not restart the Gateway or grant maintenance authority. Failed checks retain their specific diagnostic; a Gateway that is still starting keeps that outcome instead of being restarted. If command cleanup remains uncertain, the run stays open and retains its recovery artifacts instead of publishing completion or starting another repair.

Doctor repair uses the same enabled-plugin and default-check selection as ordinary Doctor lint. Opt-in checks, including the managed Codex version check, do not run during routine finalization. Explicit candidate checks still run when requested with doctor --lint --only codex/managed-app-server. The version check has a five-second deadline, terminates its process group where supported, and bounds output draining when a descendant retains a pipe. A timed-out check cannot be accepted merely because its direct child exited successfully. Nonfatal Doctor warnings appear in postUpdate.doctor.warnings; finalization reports status: "warning" and exits successfully when no other step fails. Codex runtime readiness remains owned by its plugin after restart.

Finalization (including the supervisor-facing update finalize command) records phase starts and finishes immediately on stderr and in the update run ledger. Preflight admission, config validation, config backup, and core completion-cache budgets scale with the shared SQLite database and its sidecars, with a five-minute startup allowance and conservative disk throughput. Repair Doctor and its enclosing convergence phase have no automatic wall-clock deadline. Post-plugin config and readiness checks share a budget derived from the existing shared and discovered agent database families after Doctor finishes, including WAL growth. Plugin updates retain the command owner's 20-minute allowance; missing-plugin repairs retain their installer defaults. The serial plugin phase and outer finalizer process add no competing default deadline. An explicit --timeout still overrides each phase and its child commands.

A phase deadline produces exit code 1 and JSON with status: "failed", stuckPhase, elapsedMs, error, and the existing phaseTimings array. The finalizer cancels the phase, fences further writes, and waits up to the same budget for its work to settle. Repair restores and verifies the Gateway it stopped before reporting the failure and exiting. Service custody acquisition and restoration retain their own native-operation budgets outside phase cancellation. The ledger records a warning naming the timed-out phase and budget. Preserve the phase diagnostic when reporting a stalled update.

When a fresh Doctor ran in the timed-out phase, doctorOutput includes its phase (pre-plugin or post-plugin) and separate stdout and stderr diagnostics. Each stream reports receivedBytes, lastOutputAgeMs (null when silent), and a redacted excerpt capped at 256 UTF-8 bytes. The failed phase's ledger detail and stderr retain the same excerpts before exit. Capture is limited to 64 KiB per stream; exceeding that cap replaces the text with omitted: "capture-limit". An incomplete private key or a redaction error also omits the stream text. A recent output age indicates output receipt; it does not prove that a migration advanced. Output and heartbeats do not extend the phase deadline. These diagnostics do not establish that every descendant has stopped, and must not be used as rollback authorization.

Shared CLI disposers have individual five-second deadlines. Failure diagnostics and any interactive recovery finish before the ten-second exit grace starts. If the finalizer remains alive after that grace, stderr and the ledger record active resource types and unsettled disposer names, then the process exits with its recorded outcome. A retained handle cannot withhold the supervisor's result indefinitely. Both stall diagnostics also include childProcesses: up to eight descendant processes with pid, parentPid, and an executable name (command). Arguments, environment values, and executable paths are omitted. childProcessesTruncated indicates omitted entries; childProcessInspection: "unavailable" means the process list could not be read. A null command means that process's executable name was unavailable. Inspection runs only after a stall and adds at most one second to the exit bound. Phase-failure JSON includes the same fields. Preserve these diagnostics and the phase receipts when reporting a blocked child. Completion-cache refresh remains best effort when its child can be stopped within the phase budget. A phase that exceeds its overall deadline still fails finalization.

Plugin artifacts that require capability consent are not installed without an interactive review or explicit --accept-capabilities. --yes alone does not accept capability changes, and JSON mode does not prompt. An unresolved review preserves the previous plugin payload and appears in postUpdate.plugins.warnings with a PLUGIN_CAPABILITY_CONSENT_REQUIRED outcome. When required checks pass, openclaw update can complete the core update and requested Gateway restart with status: "ok"; update repair reports status: "warning" without requesting update activation. Both commands exit successfully. This also applies when a bundled plugin moves to an external package or a missing configured plugin has no install record yet; the unreviewed replacement is not installed. Automatic repair can report a deferred replacement as a notice when a usable, enabled artifact remains installed; that retained artifact still undergoes payload validation.

If the core package has already changed, run openclaw update repair in an interactive terminal to review plugin capabilities. After reviewing the changes, automation can use openclaw update repair --accept-capabilities. Acceptance applies to each artifact's recomputed declared surface during this invocation; it does not approve future capability additions.

Skipped legacy audit recovery

In native auto mode on Linux, fs-safe handles legacy audit moves on filesystems that reject no-replace rename by publishing an exclusive hard link, then removing the old name. Doctor retains its separate compatibility publisher when the native helper is missing or disabled. Native require mode refuses unsupported moves. This preserves the original inode, including later appends from an older CLI's open file descriptor. Existing destinations are never overwritten. Doctor recovers interrupted link pairs before importing; backups capture one sanitized copy without changing either live name.

If the filesystem also rejects hard links, Doctor preserves the audit files and reports a recoverable warning with the affected filename and a command targeting that state directory. Other repairs and update finalization continue. Restore hard-link support, or stop the Gateway and all CLI writers before moving the complete state directory to a compatible filesystem, then run the reported openclaw doctor --fix command. If the directory moved, update OPENCLAW_STATE_DIR in that command. Doctor never substitutes a file copy: doing so could lose later audit appends.

When a legacy audit raw archive changed other than by append, Doctor preserves it beside itself with a .quarantined-<date>-<id> suffix. The warning names the quarantined path and explains the expected append-only growth and observed change. An empty raw archive without a checkpoint is also quarantined when its sanitized companion still contains history. Doctor keeps the sanitized records and existing SQLite rows, continues later repairs, and does not repeat the warning on subsequent runs. Quarantine does not import the changed bytes or delete the archive or backups. Quarantined raw archives remain local and are excluded from portable backups; sanitized companions and retained SQLite audit history are backed up normally.

Doctor can leave other legacy audit sources in place when a raw archive has no checkpoint and begins with ambiguous whitespace, or cannot obtain another durable raw-archive checkpoint. These conditions produce a skipped migration receipt with a warning. Other repairs continue, and update finalization can complete with warnings. An unsafe recovery failure, such as an interrupted archive that cannot be restored, still stops Doctor.

Preserve the reported source, its sanitized companion (for example, logs/config-audit.jsonl.migrated beside logs/config-audit.jsonl.migrated.raw), and any recovery journals or backups. Follow backup guidance before attempting recovery, and include the warning and archive filenames when requesting help. Do not delete or rewrite archives or checkpoints to suppress the warning.

Warnings for sources left in place repeat on later Doctor or openclaw update repair runs until the archive is resolved. Successful finalization does not mean this historical audit data was imported. There is currently no supported sanitized-only import when the raw archive is unusable: accepting the companion as a recovery source needs an explicit reconciliation procedure that preserves duplicate events, retained history, and checkpoint evidence.

update cleanup

Retire migration recovery originals and superseded original-state update captures after you have verified that the upgrade and session history work. Start with a preview, which can run while the Gateway is active:

openclaw update cleanup --dry-run
openclaw --profile work update cleanup --dry-run --json

Cleanup targets the selected profile and OPENCLAW_STATE_DIR / OPENCLAW_CONFIG_PATH overrides. It displays that state directory and does not redirect to a managed service. Confirm the displayed directory is the installation you intend to clean. --dry-run reads configuration and recovery metadata, including existing update history for startup-migration backups, without taking a maintenance lock, loading plugins, or creating state. Candidate bytes still require identity verification; historical artifacts are listed separately as requiring verification. Protected and blocked artifacts include reason codes.

Before applying, stop the Gateway for that same profile/state directory and wait for other SQLite maintenance commands to finish. Stop database readers too, including watchers that repeatedly run openclaw sessions --all-agents --json, and keep them stopped until cleanup exits. Read-only SQLite connections can create or change WAL/SHM sidecars, invalidating cleanup's destination check even when session content is unchanged. If cleanup reports Recovery destination database changed; preview cleanup again., stop those readers, preview again, and retry. Cleanup requires exclusive offline state ownership and never stops or restarts a service itself.

Cleanup permanently removes the selected rollback originals, including branches and metadata intentionally removed by a verified repair. Doctor restore cannot recreate them afterward. Keep them, or preserve an independent backup containing them, if you still need that rollback path. Current SQLite history stays in place.

openclaw update cleanup
openclaw update cleanup --yes --json

Interactive confirmation defaults to No. JSON mode never prompts or grants consent; unattended deletion requires --yes. Consent does not override ownership, file identity, or dependency checks. Applicable flags (--dry-run, --yes, and --json) work before or after cleanup; update-only flags --channel, --tag, --timeout, --no-restart, and --accept-capabilities are rejected.

Only owner-recorded recovery artifacts with complete import evidence are eligible. Unknown or unimported history, malformed inputs, trajectories, forensic corrupt databases, operator backups, and unmanifested artifacts stay protected. Old manifests are verified offline where possible; missing evidence is a reason to retain an artifact. Cleanup has no automatic expiration policy. Immutable release-retention inspection is separate from this migration-backup cleanup and remains gated on a compatible serving bridge. Its descriptor policy and release-generation inventory record ownership without deleting directories or snapshots. They do not make immutable releases eligible for update cleanup; see the immutable release-retention inventory. Doctor's <database>.pre-startup-migration-<id>.bak groups become eligible only after Doctor verifies migration completion and update history records a successful update that started later. Until then they appear as protected. Changed or unrecorded backups remain protected or blocked; cleanup never infers permission to delete from their filenames. Keep these files with your pre-upgrade backups while you still need to restore the matching database generation. Private package, command-shim, and Git runtime backups remain owned by the update transaction and are outside this migration cleanup. An interrupted entry in update history does not block cleanup of otherwise eligible migration archives.

Original-state update captures in <state-directory>.update-captures/ are listed as well, including captures retained beside the previous default state directory after migration. Each capture directory reports its logical bytes. The text summary adds an Update captures: line with their total, candidate, and protected bytes. Cleanup attributes a capture only through its run id in update history and never reads the captured payload:

  • candidate / superseded-update-capture: a sealed capture from a finished update, when a later update succeeded.
  • candidate / unsealed-update-capture: a capture that never sealed (the update reported Original update capture failed) from a finished update, when a later update succeeded. Nothing can restore from an unsealed capture.
  • protected / awaiting-later-completed-update: the capture of the latest successful update, or of any update that finished after it.
  • protected / unfinished-update-run or pending-update-recovery: the update is still running, or its recorded rollback has not settled.
  • protected / unresolved-failed-update: a sealed capture from a failed update that is neither restored nor repaired forward. It remains that update's manual recovery source.
  • protected / unmanifested-update-capture: a directory without the capture marker or without a matching update run, such as a standalone Doctor capture, and any other file in the capture directory. unreadable-update-history also keeps every capture protected.

Retiring a capture deletes its whole directory, including any candidate and prepared recovery generations inside it, under the same exclusive ownership as the rest of cleanup. An interrupted removal leaves the remainder attributed to the same run, so rerunning cleanup finishes it.

The JSON result contains stateDir, status, artifacts, and totals. Each artifact reports its path, run ids, logical bytes, outcome, and reason; update captures also carry kind: "update-capture". Totals separate candidates, verification-required, protected, blocked, and removed bytes, including update captures. Removal failures exit nonzero. Keep the recovery manifests and rerun cleanup to finish recorded interrupted work; a retry does not delete a recreated file. Removed logical bytes do not promise equivalent physical space reclamation on cloned or snapshotted filesystems. When a path cannot be inspected, its logical size comes from recorded artifact metadata when available. Cleanup records durable intent before removal and uses exclusive no-copy publication. Failures are reported; retries reconcile file operations that already completed. Manifest files are synchronized before removal; parent directories are synchronized where supported. Windows does not provide the same parent-directory durability guarantee.

Doctor restore reports intentionally disposed originals and pending cleanup explicitly. Neither update nor cleanup creates an automatic full-state backup; these recovery originals are not a full pre-upgrade backup. See Before updating: create a verified backup for backup coverage and Doctor recovery for restoring retained originals.