anthropics/skills180kwebapp-testing
Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
浏览器自动化
Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed. The full decision matrix for OOM (--novram / --cache-none / --disable-smart-memory), shared-VRAM creep on Windows (--reserve-vram N), model-switching with big text encoders (--cache-none), high-VRAM throughput (--gpu-only / --highvram), and attention-backend selection (--use-sage-attention for speed, --use-pytorch-cross-attention as the highest-quality / Z-Image-safe fallback). Also the acceleration-stack + Blackwell/RTX 5000 (sm_120) notes. Use when a graph OOMs (especially long video like LTX 2 / WAN), when the GPU spills into shared VRAM and slows to a crawl, when switching between models eats all RAM, when Z-Image produces black/garbled output under Sage, or when deciding which attention backend to launch with. Flag names verified against upstream comfy/cli_args.py; see Sources.
把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。
读取 https://funcoding.ai/skills/artokun/comfyui-mcp/comfyui-launch-flags/install.md ,按里面的步骤帮我安装这个 Skill。
CLI flags passed to main.py control ComfyUI's runtime behavior
(e.g. python main.py --reserve-vram 2 --use-sage-attention). The three that
matter most for making a graph run rather than OOM or crawl are the
VRAM strategy, the attention backend, and the cache mode. This skill
is the decision matrix for choosing them.
⚠️ Verification note (August 2026). Every flag below was checked against upstream
comfy/cli_args.pyon current master. ComfyUI adds/renames flags often — when in doubt runpython main.py --helpin the target install and prefer that over this list.--enable-triton-backend/--disable-triton-backendARE ComfyUImain.pyflags on master (they used to be documented as SwarmUI-only; that is stale).--use-ck-attentionis kitchen INT8 attention — nosageattentionwheel. A June ComfyUI checkout still pins comfy-kitchen 0.2.10 and lacks--use-ck-attention;kitchenaction:"status" reports ComfyUI-side flag support, not only the kitchen version. Usekitchen/panel_kitchento see what this GPU can actually run.
How to apply today. The MCP's
restart_comfyui(withaction: "start") currently replays the exact argv of the previous run. It does not compose fresh flags. So set these when you launch ComfyUI yourself (thepython main.py …line, arun.bat/shell alias, or the SwarmUI backend args box), and the tool will preserve them on restart. Injecting flags through the tool is a tracked follow-up.
Symptom ▶ Flag(s) to try
─────────────────────────────────────────────────────────────────────────────
CUDA out of memory, long video (LTX 2 / WAN) ▶ --novram (+ --cache-none)
OOM, still want models resident when they fit ▶ --reserve-vram N then --disable-smart-memory
GPU slows to a crawl, spills into "shared GPU ▶ --reserve-vram 2..4
memory" (Windows WDDM) mid-run
RAM blows up switching between models, or a huge ▶ --cache-none
text encoder (FLUX 2 / Mistral) won't unload
Plenty of VRAM (48GB+), want max throughput ▶ --gpu-only or --highvram
Want faster sampling on NVIDIA ▶ --use-ck-attention if kitchen INT8 is available (skip the sage wheel); else --use-sage-attention
Z-Image produces BLACK / wrong output ▶ --use-pytorch-cross-attention (NOT sage)
Sage gives black output on some models ▶ --use-pytorch-cross-attention (or fix dtype)
ROCm, kitchen present, triton ≥ 3.7 ▶ --enable-triton-backend
VRAM strategy and attention backend are each mutually exclusive groups, so
pass at most one from each. You can combine one VRAM flag + one attention flag +
one cache flag (e.g. --novram --use-sage-attention --cache-none).
| Flag | What it does | Use when |
|---|---|---|
--gpu-only | Keep everything (incl. text encoders) on GPU | 48GB+ card, single model, max speed |
--highvram | Keep models resident in VRAM after use | High-VRAM card, repeated runs of one model |
| (default) | ComfyUI's smart offload | Most setups — try this first |
--lowvram | Offload text encoders / parts to CPU | Mid card OOMing on load |
--novram | Extreme offload — minimal VRAM footprint | OOM on long video / huge models; pair with --cache-none |
--cpu | Everything on CPU (very slow) | No usable CUDA GPU only |
Modifiers (combine with the above):
--reserve-vram N reserves N GB for the OS and other apps. It is the fix for the
Windows failure mode where the GPU quietly starts using shared VRAM and
throughput collapses. Typical 2 to 4; bump to 10 for heavy video decode.--disable-smart-memory forces aggressive offload to regular RAM instead
of keeping models cached in VRAM. Reach for this when a run gets stuck or
OOMs intermittently. Slightly slower, much more reliable.--async-offload enables async weight offload streams (default on where
supported); --disable-async-offload turns it off if it misbehaves.| Flag | Notes |
|---|---|
--use-ck-attention | Comfy Kitchen INT8 attention. No sageattention wheel. Needs comfy-kitchen present and int8_attention_is_available() on this GPU. Prefer this over the sage wheel-matching install when kitchen action:"status" says INT8 is available. Restart required. |
--use-sage-attention | Quantized SageAttention kernel, ~20–40% faster sampling. Needs the sageattention package installed and version-matched — see triton-sageattention. Skip this dance when --use-ck-attention is available. |
--use-flash-attention | FlashAttention kernels. Needs flash-attn built for your torch/CUDA. |
--enable-triton-backend / --disable-triton-backend | Enable or disable the comfy-kitchen triton backend. ComfyUI master flags (not SwarmUI-only). ROCm hosts with kitchen + triton ≥ 3.7 want --enable-triton-backend. Restart required. |
--use-pytorch-cross-attention | PyTorch SDPA. Highest quality, always available, no extra deps. The safe default and the correct fallback. |
--use-split-cross-attention / --use-quad-cross-attention | Memory-optimized math attention for older/low-VRAM cards. |
Two gotchas worth memorizing:
--use-sage-attention; you get black or garbled output.
Launch Z-Image with --use-pytorch-cross-attention instead. See
z-image-txt2img.--use-pytorch-cross-attention, or (SwarmUI) set
Advanced Sampling → Preferred DType = Default (16-bit). Sage-on vs Sage-off
also produces slightly different images, so expect non-identical seeds.When a graph hard-crashes with
No module named 'sageattention'/triton: unavailable, the fix is the sdpa / no-compile fallback intriton-sageattention, not this flag.
| Flag | Effect |
|---|---|
(default --cache-ram) | Cache results under RAM pressure |
--cache-classic | Aggressive result caching |
--cache-lru N | Keep at most N node results (LRU) |
--cache-none | Cache nothing — re-executes every node; lowest RAM/VRAM. Essential when switching between dual models or when a giant text encoder (FLUX 2's Mistral) must fully unload. |
--fast enables experimental, potentially quality-degrading
optimizations. Accepts specific PerformanceFeature values:
fp16_accumulation, fp8_matrix_mult, cublas_ops, autotune. Bare --fast
turns them all on. Test output quality before committing to it.--fp8_e4m3fn-unet, --fp16-unet, --bf16-unet, --fp32-unet, …) for
forcing a compute precision. Usually the model or loader picks the right one, so
only reach for these to work around a specific dtype error.Long video OOM (LTX 2 / WAN, 24GB): --novram --cache-none
(add --disable-smart-memory if it stalls)
Windows shared-VRAM creep: --reserve-vram 3
FLUX 2 / huge text-encoder swaps: --cache-none
High-VRAM throughput (48GB+): --gpu-only (or --highvram)
Fast NVIDIA sampling (most models): --use-ck-attention (if kitchen INT8 is available)
--use-sage-attention (otherwise; needs the wheel)
Z-Image (any): --use-pytorch-cross-attention
ROCm + kitchen + triton ≥ 3.7: --enable-triton-backend
Cross-refs: video OOM specifics in
ltxv2-video / wan-t2v-video;
per-model VRAM math in troubleshooting and
model-compatibility.
The attention/compile accelerators are version-locked to your exact
torch + CUDA + Python. A mismatched wheel doesn't just fail to import; it can
break the torch install. A known-good, mutually-compatible stack for late-2025 /
2026 NVIDIA (including Blackwell / RTX 5000, sm_120) looks like:
| Component | Role | Notes |
|---|---|---|
| Torch + CUDA | base | e.g. Torch 2.9.x on CUDA 12.8/13; use the wheel index matching your driver |
| Triton | torch.compile / inductor | Windows: triton-windows (woct0rdho) |
| SageAttention | --use-sage-attention | wheel matched to torch/CUDA/python |
| FlashAttention | --use-flash-attention | built per torch/CUDA/python |
| xFormers | memory-efficient attention | optional |
| InsightFace | FaceID / IP-Adapter / ReActor | onnxruntime-gpu alongside |
Operational facts worth carrying:
TORCH_CUDA_ARCH_LIST=7.5;8.0;8.6;8.9;9.0;10.0;12.0+PTX spans RTX 20xx→50xx
and datacenter (A100/H100/B200). +PTX lets newer archs JIT.~/.triton / %USERPROFILE%\.triton and temp)
when you hit stale-kernel Triton errors after an upgrade.uv pip install over pip for the venv. Resolves and downloads are
dramatically faster. install_comfyui already supports this via preferUv.troubleshooting.--use-ck-attention, --enable-triton-backend, --disable-triton-backend, --fast); hardware gates in comfy/model_management.py (supports_fp8_compute SM ≥ 8.9, supports_nvfp4_compute / supports_mxfp8_compute SM ≥ 10.0); kitchen backends in the comfy-kitchen README https://github.com/Comfy-Org/comfy-kitchen--enable-triton-backend is retracted as of ComfyUI master.
anthropics/skills180kToolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
浏览器自动化
addyosmani/agent-skills103kTests in real browsers via Chrome DevTools MCP. Use when building or debugging anything that runs in a browser. Use when you need to inspect the DOM, capture console errors, analyze network requests, profile performance, or verify visual output with real runtime data. Requires the chrome-devtools MCP server to be configured.
浏览器自动化
ComposioHQ/awesome-claude-skills77kToolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
浏览器自动化
code-yeongyu/oh-my-openagent70kDrives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch.
浏览器自动化
shanraisshan/claude-code-best-practice67kBrowser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking screenshots, extracting data, testing web apps, or automating any browser task. Triggers include requests to "open a website", "fill out a form", "click a button", "take a screenshot", "scrape data from a page", "test this web app", "login to a site", "automate browser actions", or any task requiring programmatic web interaction.
浏览器自动化
CherryHQ/cherry-studio52kRun Cherry Studio critical-path system regression tasks through the repository-owned Playwright E2E workflow. Use for full regression, release acceptance, development-branch system validation, or a named cherry-regression-test task on GitHub-hosted macOS and Windows runners.
浏览器自动化