# Sub-agent concurrency, recovery, and stopping

> The subagent queue lane, restart recovery, stop scope, and the standing limitations

- 网址：https://funcoding.ai/agents/openclaw/tools/subagents/operations/
- 来源：OpenClaw 官方文档原文（英文），MIT 许可，同步于 2026-10-11
- 官方原文：https://docs.openclaw.ai/zh-CN/tools/subagents/operations

---
## Concurrency

Each spawning session has its own in-process queue for ordinary sub-agents. The setting
`agents.defaults.subagents.maxConcurrent` limits concurrent child runs for that
session (default `8`). Independent sessions have independent budgets, so one
busy session does not consume another session's sub-agent slots. A nested
orchestrator's children use that orchestrator's budget, not its parent's.

The budget belongs to the child's immediate spawning/controller session.
Changing where a child's completion is delivered does not move its execution
to another session's budget. Accepted runs above the execution limit queue until
a slot is available.

[Swarm](https://funcoding.ai/agents/openclaw/tools/swarm/) collector children (`collect: true`) instead use a dedicated
`subagent:swarm:<schedulerGroupKey>` lane. Its cap is the group's resolved
`tools.swarm.maxConcurrent` (default `32`), independent of the parent's ordinary
`subagent:<immediate session>` lane. An ordinary spawn from the same parent can
start while its Swarm lane is full. Ordinary children spawned by a collector use
that collector's own session lane. Each running collector costs one model stream
and one Code Mode worker isolate; tune the Swarm cap for the Gateway's resources.

`maxChildrenPerAgent` is a separate admission limit on active children per
session (default `5`); increasing execution concurrency does not raise that
limit. Collector admission instead uses Swarm's `maxChildrenPerGroup` (default
`50`) and `maxTotalPerGroup` (default `200`); raising execution concurrency does
not raise either group limit. [Codex-native subagents](https://funcoding.ai/agents/openclaw/plugins/codex-harness/) use Codex's own scheduler
and limits independently of these OpenClaw queues.

Suspended completion deliveries do not block new work. Native subagents, ACP
sessions, and visible sessions retain their normal active-run limits and
authorization checks independently of the delivery backlog.

OpenClaw warns when the delivery backlog reaches 25. Within a Gateway process,
unchanged backlog counts do not repeat the warning every sweep. A count change
at or above 25, or a return to that threshold after recovery, produces a new
warning. The backlog size does not discard results or change their retention.

A suspended delivery expires after seven days. The warning records its run and
session identifiers, original delivery error, and suspension time. Expiry ends
automatic delivery; it does not turn a successful child execution into a failure.
For a retained run, `/subagents info <runId>` shows the captured result, discarded
delivery status, and original error, even after the run leaves the recent list.
Inspect that result or available child session history and ask the requester to
continue. Cleanup and retention still apply: `cleanup: "delete"` can remove the
run, and hidden session cleanup can remove its transcript. Expiry does not promise
that deleted results remain recoverable.

## Liveness and recovery

OpenClaw does not treat `endedAt` absence as permanent proof that a
sub-agent is still alive. When the reading process can verify a current
execution owner or an exact queued collector reservation, an unended run keeps
counting regardless of age. Persisted metadata alone does not establish that
ownership in another process. Other unended runs stop counting as active/pending
after the stale-run window
(2 hours, or the configured run timeout plus a short grace period,
whichever is longer). These retained counts govern `/subagents list`,
status summaries, descendant completion gating, and per-session concurrency
checks; they are not proof that an executor is live.

During a graceful restart or Gateway suspension, lifecycle events from
already-admitted runs remain tracked through restart preservation and completion
processing. An already-admitted replacement run can finish refreshing a deferred
child result before shutdown. The refresh remains tracked until capture and
persistence finish; it does not admit a new run.

Completion waits retire with their original registry database during shutdown.
Late cleanup recovery leaves retained state for the next Gateway instead of
retrying against a closed store. Failed writes with an unknown outcome still
report an error and remain fenced until recovery reads the canonical state.

After a Gateway restart, the parent owns continuation of the user's task.
Interrupted sub-agents are finalized through their normal completion path instead
of automatically relaunched. Their results tell the parent that execution was
interrupted and that partially completed actions need checking. A parent waiting
for its child batch receives the settled results, including children that finished
before the restart, and decides what work remains.

An already-admitted completion turn for a yielded nested requester is different
from an orphaned child launch. Its frozen child-result batch retains the exact
saved continuation across restart. Registry recovery waits for that owner instead
of reporting interruption while the same continuation is being replayed. This
does not authorize automatic relaunch of unrelated interrupted child work.

Recovery retires superseded requester-transfer generations before restoring the
remaining claim. The current generation keeps its completion custody; obsolete
transfers do not retry indefinitely. Transient persistence failures still retry.

Recovery handles both sessions marked `abortedLastRun: true` and hard kills that
prevented the shutdown marker from being written. For a hard kill, the child
session must still identify the exact run from the retired Gateway, with no newer
run or admitted work owning that session. Live executions and queued collectors
keep their existing owners. Orphaned runs settle their native execution state
before cleanup, so retained child sessions do not leave phantom running activity.
If settlement fails, completion remains available for retry.

Recovery remembers live session ownership without retaining saved prompt snapshots.
Unchanged owned rows skip session reads; ownership release or a published session
change requests another sweep within one second. Unavailable recovery reads retry
per row with exponential backoff from one second to one minute. Registry lifecycle
changes and session publications clear that backoff. Persisted rows without a live
owner still require a fresh read before recovery can change their state.

Startup session maintenance reports retained run/task owners in one informational
summary. Those rows remain with registry recovery; the session-only orphan repair
does not compete for their ownership. Ownership changes during a repair and failed
ownership checks still produce warnings.

The parent can inspect a retained child transcript and use `sessions_send` to
continue that session, or spawn a replacement after confirming the old execution
has stopped. Reusing a child restores its conversation context; it does not replay
an interrupted command. Existing cleanup and retention settings still apply.

When a detached cleanup attempt logs `subagent cleanup finalize failed`, its
retries use bounded backoff in the current Gateway process. If those retries are
exhausted, the run remains recorded with incomplete cleanup; inspect the warning
to identify the failing operation. Unrelated
sub-agent completions do not restart failed cleanup or reset its retry budget.
Descendant completion still wakes the current requester ancestors waiting on that
work. These cleanup retries are separate from [completion delivery](https://funcoding.ai/agents/openclaw/tools/subagents/announce/).

When required registration has an unknown outcome or retained registry state
forbids deleting the child session, a Gateway-hosted ordinary spawn's error keeps
the child's session and run identifiers. If the first cancellation attempt also
fails, the error reports unconfirmed termination and whether Gateway cleanup was
scheduled. The Gateway retains the child's admission slot while retrying and
rechecks the exact run owner before each attempt. If database admission retires
while that child still runs, cancellation stops but its slot stays reserved until
the child controller retires or Gateway shutdown takes over. The child session
stays intact. Provisional session rollback, collector FIFO cleanup, and local
embedded cleanup remain joined. Inspect the retained child before retrying the
spawn.

<div class="callout callout-note">

If a sub-agent spawn fails with Gateway `PAIRING_REQUIRED` /
`scope-upgrade`, check the RPC caller before editing pairing state.
Internal `sessions_spawn` coordination dispatches in process when the
caller is already running inside the gateway request context, so it does
not open a loopback WebSocket or depend on the CLI's paired-device scope
baseline. Callers outside the gateway process still use the WebSocket
fallback as `client.id: "gateway-client"` with `client.mode: "backend"`
over direct loopback shared-token/password auth. Remote callers, explicit
`deviceIdentity`, explicit device-token paths, and browser/node clients
still need normal device approval for scope upgrades.

</div>

## Stopping

An explicit Stop targeting a parent run cancels the children associated with that
run and their descendants, including ordinary sub-agents and [Swarm](https://funcoding.ai/agents/openclaw/tools/swarm/)
collectors. Successful cancellation keeps selected queued collectors from
starting while running children stop. Exact-run cancellation does not cancel
unrelated turns or clear unrelated session-wide queues.

For raw child session keys such as `global`, registration records the known
owning agent as `childAgentId`. Child cancellation derives that agent's session
store from current configuration, just as it does for agent-qualified keys,
which need no recorded binding. Clearing the child's queued follow-ups and
commands stays within that agent, even when another agent uses the same raw
session key. Legacy runs without a recorded owner use current configuration to
resolve one agent and clear only that agent's queues. The optional binding stays
in `payload_json` when an older build rewrites the run.

Owner-aware child lookups keep watched follow-ups, steering, run generations,
completion transcripts, and timeout reconciliation separate when agents share a
raw key. Terminal events and session timing updates use that same child owner.
Rows without a recorded owner and callers without an explicit owner retain their
existing key-only lookup behavior.

Stop also retires pending completion continuations for the selected work, even
when a child has already finished. Cancelling a completion turn retires its
matching child batch, so automatic delivery retries cannot start it again under
a new run ID. Captured child results and their execution outcomes remain intact;
you can inspect them or send a new instruction afterward.

A parent Stop also drains ordinary commands retained by its selected native
children, even after their model execution has completed. Completed child results
remain completed; command cleanup does not add canceled model runs to the result.
Commands deliberately started with `background: true` remain independent.
A live command continuation retains the original request's child ownership when
it runs in another session. Stop selects only that request's bound native children,
leaving newer human work and replacement child generations alone.
A fresh session Stop can still select those children after the continuation has
completed and its notification has been consumed. It uses the current original
session's authority and generation; a historical exact-run Stop does not regain
authority after that request's active, process, and event ownership has ended.

For Gateway callers, `chat.abort` with a `runId` uses this exact-parent scope.
`sessions.abort` with a `runId` also targets that run. When it resolves a recovered
native run without a chat controller, it cancels children only if the captured
active parent accepts Stop; a declined or no-active-run result, including an
already-finalizing parent, leaves those children alone.

Sending `/stop` in the requester chat has broader scope: it aborts requester
session work, clears its queues, and cancels its active child tree. Session-wide
`sessions.abort` also requests descendant cancellation; clearing queued follow-ups
requires `clearQueued: true`. Ordinary `chat.abort` without a `runId` does not
cascade to children. These operations retain their normal authorization checks.

A typed `/stop` sent through `chat.send` honors `expectedLeafEntryId` and, when
that branch check is present, `sessionId`. If the check fails during descendant
cancellation, the Gateway refuses further cancellation and reports
`active-leaf-changed`. Cancellation already accepted by a child still settles.
The cancellation result waits for pending child-session metadata writes before
checking that the original session still owns the outcome.

Incomplete cancellation is reported as an error, not a clean success. `/stop`
reports actual stopped and failed child counts. A committed child cancellation
remains in the stopped count if later cleanup fails or the parent is replaced.
The cleanup error remains visible; a replacement run requires its own Stop
request. Inspect the remaining native subagent runs with `subagents` and retry
their cancellation; request acknowledgment does not mean all runtime cleanup is
instantaneous.

Accepted children remain independent after ordinary parent completion, yield, or
timeout. Those events do not automatically cancel them.

## Limitations

- Direct announce attempts are best-effort, but admitted session-queued completion handoffs and their native subagent records survive gateway restarts in the shared SQLite state database.
- Sub-agents still share the same Gateway process resources; concurrency caps apply per spawning session or Swarm group, not to total Gateway concurrency.
- `sessions_spawn` returns `{ status: "accepted", runId, childSessionKey }` when startup is accepted, without waiting for the child task to finish. Cloud-worker spawns can wait for provisioning before returning this receipt.
- Sub-agent context only injects `AGENTS.md` (no `SOUL.md`, `IDENTITY.md`, `USER.md`, `MEMORY.md`, or `BOOTSTRAP.md`). Its `## Tools` section carries environment-specific notes. Codex-native subagents follow the same boundary through native `AGENTS.md` discovery, while parent-only persona, identity, and user files are injected as turn-scoped collaboration instructions so children do not clone them.
- Recursive spawning is enabled through depth `5` by default. Set `maxSpawnDepth` from `1` through `5` to lower the boundary.
- `maxChildrenPerAgent` caps active children per session (default `5`, range `1-20`).
