跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

hyperpod-nccl

Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures, EFA TCP fallback, /dev/shm or memlock issues, NCCL version mismatch across pods, container OOM / exit-137 / OOMKilled, GPU OOM (CUDA out of memory), CrashLoopBackOff / Pending pods, MASTER_ADDR DNS, NetworkPolicy blocking. Not for single-node hardware faults (→ hyperpod-node-debugger § G) or cluster-creation EFA / SSM failures (→ hyperpod-cluster-debugger § A / § F).

DevOps 与云915plugins/sagemaker-ai/skills/hyperpod-nccl/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/awslabs/agent-plugins/hyperpod-nccl/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

HyperPod NCCL Debugger

Operating policy. Run read-only diagnostics yourself. Never run a command that changes cluster, node, or workload state — present each one as a Suggested command (run this yourself) block and wait for the customer. Destructive order: investigate → reboot → replace (replace destroys root + secondary volumes; not supported on Slurm controller nodes). Never discard training state on speculation.

Diagnose NCCL failures on SageMaker HyperPod (EKS and Slurm). scripts/nccl-diagnose.sh reads state via AWS APIs, kubectl, and SSM, then prints each issue as [FAIL] ... → references/<file>.md § <section>. Read-only.

Signal sourcing: list-cluster-events carries infrastructure-level state only (lifecycle, bootstrap, EFA health check, capacity, replacement, reboot, AMI rollback). It does not carry NCCL timeouts, GPU XID/ECC, or per-pod training signals — those come from pod logs, CloudWatch training streams, on-node SSM probes, and NCCL env audit. "No events" on a training-time NCCL issue is expected, not a clean bill of health.


Workflow

  1. Collect cluster name, region, namespace/job (EKS), exact NCCL error string.
  2. Run the diagnostic (always — the output drives everything else).
  3. For every [FAIL] line, Read the referenced section.
  4. Present finding, root cause, and the Suggested-command block with concrete values (instance IDs, SG IDs, namespaces) filled in from the script output. Wait for customer approval.
  5. Re-run the diagnostic to confirm.

If a finding has no matching section, report it as a bug — do not invent a fix.

Step 1: Authenticate kubectl (EKS)

EKS_ARN=$(aws sagemaker describe-cluster --cluster-name <HYPERPOD-NAME> --region <REGION> \
  --query 'Orchestrator.Eks.ClusterArn' --output text)
EKS_NAME=$(echo "$EKS_ARN" | awk -F'/' '{print $NF}')
aws eks update-kubeconfig --name "$EKS_NAME" --region <REGION>
kubectl get nodes

Step 2: Run the diagnostic

# Basic:
bash scripts/nccl-diagnose.sh --cluster <HYPERPOD-NAME> --region <REGION>

# Scope to an EKS job/namespace:
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --namespace <NS> --job <JOB>

# Force orchestrator:
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --orchestrator slurm

# Larger hardware sample (default 3):
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --sample-nodes 10

# Specific node only:
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --node i-0abc123def456

Tags: [PASS] · [FAIL] (counted in Issues Found, has reference pointer) · [WARN] · [INFO]. Priorities: P0 blocks training · P1 degraded · P2 informational.


Remediation index

Each [FAIL] line in the script already points directly at the right section. This table is a lookup for manual triage.

FindingSection
SG missing inbound/outbound self-referenceoperations.md § 8
Blocking NetworkPolicy / allow-all missingoperations.md § 8
Slurm node DOWN / DRAINING / RemoveIPCoperations.md § 7
GPU XID / SYSTEM_ERROR / hardware faulthyperpod-node-debugger § F / § G
GPU row-remap / DCGM Fail / silent NaNshyperpod-node-debugger § G.1.a/b
NCCL timeout / rendezvous / stragglerdebugging-guide.md § 1
EFA configuration / not useddebugging-guide.md § 6
EFA TCP fallback (NET/OFI Using TCP)debugging-guide.md § 13
NCCL version mismatch across podsdebugging-guide.md § 10
Container OOM (pod killed, exit 137)debugging-guide.md § 4
GPU OOM (CUDA out of memory)debugging-guide.md § 11
RDMA memlock / /dev/shm too smalldebugging-guide.md § 17
MASTER_ADDR DNS / headless Servicedebugging-guide.md § 12
NVLS / PXN / topology tuningdebugging-guide.md § 19
Any NCCL / EFA / rendezvous log patternerror-patterns-quick-ref.md
Performance / nccl-tests / bandwidthperformance-testing.md

Prerequisites

  • aws CLI v2.13+ authenticated (aws sts get-caller-identity)
  • jq, python3, bash 4.2+
  • unbuffer (from the expect package: yum install expect / apt install expect)
  • kubectl authenticated to the EKS cluster (K8s checks skipped if absent)
  • session-manager-plugin for on-node hardware checks

Defaults

  • Region — required: pass --region or set $AWS_DEFAULT_REGION.
  • Orchestrator — auto-detected; override with --orchestrator eks|slurm.
  • Namespace / job (EKS) — all namespaces; scope with --namespace <NS> --job <JOB>.
  • Hardware sampling — 3 nodes over SSM (capped at 50). --node <ID> for a specific node. Node probes run serially (180 s per node): --sample-nodes 10 can take ~30 min.
  • CloudWatch window — last 2 hours.
  • Colors — auto-disabled on non-TTY or TERM=dumb.

Error handling

FailureScriptTell the customer
aws sts get-caller-identity failsExit 1 with the AWS error"Fix AWS credentials and rerun."
describe-cluster AccessDeniedWarn, add Missing IAM for sagemaker:DescribeCluster"Grant sagemaker:DescribeCluster (operations.md § 2)."
Cluster not foundExit 1 after listing region's clusters"Confirm HyperPod cluster name and region."
kubectl absent / unauthenticatedWarn, skip K8s checks"aws eks update-kubeconfig --name <EKS> --region <R>."
SSM plugin absentWarn, skip on-node hardware checks"Install session-manager-plugin."
SSM times out (180s)Partial output, mark node unreachable"Rerun with --node <ID> --sample-nodes 1; check SSM agent on the node."
CloudWatch log group not foundSkip CloudWatch scan"Enable CloudWatch on the cluster (operations.md § 4)."
Cluster events API throttledWarn, continue with partial data"Rerun later — script is idempotent."

Exit codes: 0 diagnostic complete · 1 fatal prerequisite missing or cluster unreachable.

IAM permissions

Full policy + RBAC in operations.md § 2. SSM on HyperPod uses start-session against sagemaker-cluster:<cluster-id>_<group>-<iid> targets — grant ssm:StartSession / ssm:TerminateSession, not ssm:SendCommand.

Scale strategy

ScopeMethodCoverage
All nodessagemaker:ListClusterNodes (paginated)100% nodes
All K8s objectskubectl100% pods/nodes/policies
HardwareSSM --sample-nodes N (default 3)Sampled
Node logsCloudWatch100% nodes

Large clusters: the PyTorch NCCL backend defaults to a 10-minute collective-op timeout (per the PyTorch distributed docs). Large clusters routinely exceed that on first rendezvous; raise it via torch.distributed.init_process_group(timeout=timedelta(seconds=<N>)). HyperPod support has also observed NCCL topology-graph-search hangs on 256+ node clusters when memlock is unlimited; using a large fixed memlock (e.g. 8388608) in pod securityContext or /etc/security/limits.conf has cleared these in field cases. This memlock pattern is a field observation, not AWS- or NCCL-documented behavior.

For FSDP, DeepSpeed, or Megatron-LM tuning: debugging-guide.md § 18.

Skill delegation

NeedUse
Cluster creation / deployment failureshyperpod-cluster-debugger (§ A / B / C / H + --validate)
Post-deployment cluster-wide managementhyperpod-cluster-debugger
Per-node issues (disk, lifecycle, hardware)hyperpod-node-debugger
Trainium/Inferentia collective-comm (AWS Neuron Collectives, not NCCL)hyperpod-node-debugger § G.2
Shell on nodeshyperpod-ssm
Version comparison across nodeshyperpod-version-checker
Diagnostic bundle for AWS Supporthyperpod-issue-report
MFU / performance degradationhyperpod-mfu-debugger

Escalate to AWS Support

Escalate when:

  1. All SG rules correct, EFA verified on-node, but NCCL still times out.
  2. Hardware checks pass on all nodes but AllReduce still hangs.
  3. Issues Found: 0 but training still fails.
  4. GPU XID errors persist after node replacement.
  5. Collective-op timeout raised and memlock workaround applied but large-cluster rendezvous still hangs.

Before opening the case

# 1. Cluster identity + status
aws sagemaker describe-cluster --cluster-name <C> --region <R>

# 2. Full NCCL diagnostic (sample more nodes for escalation)
bash scripts/nccl-diagnose.sh --cluster <C> --region <R> --sample-nodes 10 > nccl-diag.txt

# 3. Per-node log/config bundle to S3 (delegates to hyperpod-issue-report)
#    See skills/hyperpod-issue-report/SKILL.md for the exact invocation.

Include in the case

  • Cluster name + ARN and AWS region
  • Orchestrator (EKS or Slurm) and EKS cluster name / Slurm controller node
  • Timestamp window (UTC start / end) of the failure
  • Exact NCCL / libfabric error strings (copy verbatim from pod logs or journalctl)
  • Affected instance IDs / node names / pod names / namespace / job name
  • nccl-diag.txt from step 2 above
  • S3 URI of the hyperpod-issue-report bundle from step 3
  • NCCL env vars in effect (printenv | grep -E '^NCCL|^FI_|^TORCH_' from one pod)

References

相似的 Skill

shipping-and-launch
addyosmani/agent-skills103k

shipping-and-launch

Prepares production launches. Use when preparing to deploy to production, or when asking what needs to be in place before shipping. Use when you need a pre-launch checklist, when setting up monitoring, when planning a staged rollout, or when you need a rollback strategy.

DevOps 与云

publish
code-yeongyu/oh-my-openagent70k

publish

Publish oh-my-opencode to npm by triggering the GitHub Actions publish workflow and verifying its artifacts. Ship-only: never runs pre-publish-review or re-reviews merged code unless the user explicitly asks. Argument: <patch|minor|major|explicit-semver>. Triggers: publish, release, deploy, npm publish.

DevOps 与云

acceptance-orchestrator
sickn33/agentic-awesome-skills47k

acceptance-orchestrator

Use when a coding task should be driven end-to-end from issue intake through implementation, review, deployment, and acceptance verification with minimal human re-intervention.

DevOps 与云

event-store-design
wshobson/agents40k

event-store-design

Design and implement event stores for event-sourced systems. Use when building event sourcing infrastructure, choosing event store technologies, or implementing event persistence patterns.

DevOps 与云

arize-ai-provider-integration
github/awesome-copilot40k

arize-ai-provider-integration

Creates, reads, updates, and deletes Arize AI integrations that store LLM provider credentials used by evaluators and other Arize features. Supports any LLM provider (e.g. OpenAI, Anthropic, Azure OpenAI, AWS Bedrock, Vertex AI, Gemini, NVIDIA NIM). Use when the user mentions AI integration, LLM provider credentials, create integration, list integrations, update credentials, delete integration, or connecting an LLM provider to Arize.

DevOps 与云

appinsights-instrumentation
github/awesome-copilot40k

appinsights-instrumentation

Instrument a webapp to send useful telemetry data to Azure App Insights

DevOps 与云