anthropics/skills180kwebapp-testing
Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
浏览器自动化
Turn a user journey you just clicked through into a saved regression check that re-runs deterministically, with no model in the loop and no test code to write. Use when you have driven the same flow twice, when the user wants regression coverage without a Playwright suite, when a refactor needs proving against every existing journey, or when re-verifying by hand is costing a full drive every time.
把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。
读取 https://funcoding.ai/skills/reticlehq/reticle/replay-user-flows/install.md ,按里面的步骤帮我安装这个 Skill。
Exploring an app to find a journey is the expensive part, and re-driving it with a model pays that cost again on every change. Reticle flows pay it once: the journey is saved with semantic anchors and replayed deterministically afterwards.
Needs Reticle wired in the project. Not there? RETICLE_INSTALL_SOURCE=npx_skill npx @reticlehq/server@latest init, then the install-and-verify skill.
You do not have to ask. A drive is saved as a flow automatically when the session ends, written to .reticle/flows/. Commit it: any agent on the repo can then replay it.
What decides whether it is worth committing is how you drove it. A step keeps a consequence only when you declared one, so drive the golden path like this:
reticle_act_and_wait({ sessionId, ref, action: "click", until: { kind: "signal", name: "task:created" } })
That step replays as a test. A bare reticle_act saves a click with nothing to prove, and replays green through any regression.
To name a flow deliberately rather than take the automatic one, reticle_record and reticle_flow_save do it through reticle_run { tool, args }. Neither is advertised, and neither needs to be: reticle_run is on the default surface and dispatches to any registered tool by name. RETICLE_ADVERTISE_ALL_TOOLS=1 advertises them outright instead, which suits a suite that calls by name rather than a running agent.
Annotate the business outcome, not just the clicks, so a replay proves the journey achieved something. Extended surface, like the two above:
reticle_run({ tool: "reticle_annotate", sessionId, args: { flow: "create-task", kind: "intent", text: "create a task and see it in the list" } })
reticle_run({ tool: "reticle_annotate", sessionId, args: { flow: "create-task", kind: "success-state", signal: "task:created" } })
You do not need to add data-testid first. A step whose element has no testid is anchored on its component and source location automatically, and a testid-preserving refactor still replays green.
reticle_verify({ sessionId, action: "change", files: ["src/tasks/TaskList.tsx"] })
That replays the flows covering those files. To replay one named flow directly, reticle_flow_replay is reached through reticle_run.
Three statuses, and the failures are legible rather than blind:
| status | means | next |
|---|---|---|
ok | every anchor resolved, every expectation held | done |
drift | an anchor missed: a renamed testid, a signal that never fired | read decision.nextAction; it names the file:line and the closest surviving anchor |
error | the flow file is missing or invalid, or a step failed at runtime | fix from the error envelope's failed step |
On drift, reticle_verify { action: "heal" } proposes the nearest-match rebind so flows do not rot. Apply it when the rename was intentional; treat it as a finding when it was not.
reticle_verify({ sessionId, action: "flows" })
// → { status, total, passed, failed, failures: [{ flow, verdict, whatChanged, whereInSource, nextAction }] }
One call, every saved flow, no model per flow. Only failures carry detail, so a green suite is cheap to check. Build → flow_verify → fix from each nextAction → repeat is the regression loop, and it is the point of recording in the first place.
On a large suite, replaying everything after a one-file edit is waste. Hand it the diff instead:
reticle_verify({ sessionId, action: "change", since: "HEAD~1" })
It works out which saved flows cover the files you edited and replays only those. Give it a git ref or the file list. Use this in the inner loop and flow_verify before you ship: the narrow one is fast, the whole one is the guarantee.
reticle_run({ tool: "reticle_domain", sessionId }) // not advertised, one hop away
// → { flowCount, coverage: { asserted, presenceOnly, assertionFree }, gaps: { declaredUntestedSignals, … } }
A recorded flow that asserts nothing replays green through any regression: it proves the clicks still resolve, not that the app still works. Check this after a recording session: anything landing in assertionFree needs an annotate pass with a success-state, or it is decoration.
A journey you will run once is cheaper to drive with reticle_act_and_wait and forget. Record the flows that define the product (the ones a regression in would be a bad day) and leave exploratory drives unsaved. A suite of forty half-meant flows costs more attention than it returns.
A replay reports what happened. drift is not a pass, and healing a flow to make it green when the app genuinely broke is the one thing that makes the whole suite worthless. If the rename was not intentional, the drift is the finding: report it with the whereInSource pointer.
Full flow reference, one page: curl https://docs.reticle.sh/flows.md. Index of everything: curl https://docs.reticle.sh/llms.txt.
anthropics/skills180kToolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
浏览器自动化
addyosmani/agent-skills103kTests in real browsers via Chrome DevTools MCP. Use when building or debugging anything that runs in a browser. Use when you need to inspect the DOM, capture console errors, analyze network requests, profile performance, or verify visual output with real runtime data. Requires the chrome-devtools MCP server to be configured.
浏览器自动化
ComposioHQ/awesome-claude-skills77kToolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
浏览器自动化
code-yeongyu/oh-my-openagent70kDrives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch.
浏览器自动化
shanraisshan/claude-code-best-practice67kBrowser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking screenshots, extracting data, testing web apps, or automating any browser task. Triggers include requests to "open a website", "fill out a form", "click a button", "take a screenshot", "scrape data from a page", "test this web app", "login to a site", "automate browser actions", or any task requiring programmatic web interaction.
浏览器自动化
CherryHQ/cherry-studio52kRun Cherry Studio critical-path system regression tasks through the repository-owned Playwright E2E workflow. Use for full regression, release acceptance, development-branch system validation, or a named cherry-regression-test task on GitHub-hosted macOS and Windows runners.
浏览器自动化