Skip to content
FunCoding

Search

Search docs, Skills and MCP

agentic-tdd

Test-driven development for behaviour a unit test cannot reach, by writing the expectation against the running app before writing the code. Declare the consequence first, watch it fail, implement, watch it pass. Use when building a user-facing feature, when the user asks for TDD on UI or full-stack work, when a unit test cannot express the outcome that matters, or when you want a red-green loop that runs against the real app instead of mocks.

浏览器自动化1.2kskills/agentic-tdd/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/reticlehq/reticle/agentic-tdd/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Red, green, refactor: against the running app

Unit tests drive the units. They cannot say "clicking Deploy posts to /api/deploy, moves the store to deploying, and shows the banner": that outcome only exists when the whole app runs. So the loop stalls exactly where the interesting bugs are, and the agent falls back to writing code and hoping.

This runs the same discipline one level up, using Reticle to drive the real app. Not installed? RETICLE_INSTALL_SOURCE=npx_skill npx @reticlehq/server@latest init, then the install-and-verify skill.

Why this is TDD and not just testing afterwards

The whole value of test-first is that the oracle is written while you still do not know the answer. An expectation written after seeing the result can always be adjusted into agreeing with whatever happened, and an agent is especially good at that adjustment. It will find a reading of the output under which the code it just wrote is correct.

reticle_act_and_wait({ ref, action, until }) enforces the order structurally: until is an argument to the action, so the consequence is named before the action runs. That is the red-green loop, made unfakeable.

1. RED: write the expectation, watch it fail

Before you write the feature, state what the app must do:

reticle_act_and_wait({ sessionId, ref, action: "click", until: { kind: "allOf", predicates: [
  { kind: "net",     method: "POST", urlContains: "/api/deploy", status: 200 },
  { kind: "signal",  name: "deploy:started" },
  { kind: "element", query: { testid: "deploy-banner" } },
  { kind: "console", level: "error", absent: true },
]}})

You want verified: "no" here. A red you did not see is a test you cannot trust. If this comes back yes before you have written anything, the expectation is not specific enough to the change. Tighten it until it fails for the right reason.

verified: "unknown" is not a red. It means Reticle could not tell, so the loop has no signal at all. Fix that before writing code, usually by naming a consequence the app can actually produce.

2. GREEN: implement until the same call passes

Write the smallest change that makes it hold, then re-run the same call, unchanged. That last word is the discipline: editing the predicate to match what you built converts TDD into narration. If the assertion has to change, say out loud why the original expectation was wrong.

Prefer re-asserting over re-driving when the verdict was unknown / unsettled: reticle_assert({ predicate, since, timeout_ms: 8000 }). Re-driving repeats a side effect that already happened.

3. REFACTOR: the expectation is the safety net

Now change the implementation freely and re-run. Predicates are bound to behaviour (a request, a signal, a state path), not to markup, so a refactor that preserves behaviour stays green while a DOM-shaped test would go red for no reason.

4. Keep the loop for the next change

A journey worth writing test-first is a journey worth re-running forever. Save it once:

reticle_run({ tool: "reticle_flow_save", sessionId, args: { flowName: "deploy" } })
reticle_run({ tool: "reticle_verify", sessionId, args: { action: "flows" } })   // every saved flow, no model per flow

That turns the red-green loop into a regression suite you never hand-wrote.

What to assert on, in order of strength

  1. A signal the app fires itself ({ kind: "signal" }): the app declaring success in its own vocabulary. Strongest available.
  2. State (reticle_look { action: "state" }): what the app believes. Catches a UI that moved while the store did not.
  3. Network: the request, method and status. Catches a mock standing in for the real thing.
  4. An element appearing: necessary, never sufficient. Anything can render.
  5. Absence of console errors: always include it, never rely on it alone. Absence-only predicates pass on a control wired to nothing.

Honesty

Never weaken a check to turn a verdict green. In this loop that is not a small sin: it is the loop running backwards, and it produces a green suite over a feature that does not work.


Predicate reference: curl https://docs.reticle.sh/predicates.md. Everything else: curl https://docs.reticle.sh/llms.txt.

Similar Skills

webapp-testing
anthropics/skills180k

webapp-testing

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

Browser automation

browser-testing-with-devtools
addyosmani/agent-skills103k

browser-testing-with-devtools

Tests in real browsers via Chrome DevTools MCP. Use when building or debugging anything that runs in a browser. Use when you need to inspect the DOM, capture console errors, analyze network requests, profile performance, or verify visual output with real runtime data. Requires the chrome-devtools MCP server to be configured.

Browser automation

webapp-testing
ComposioHQ/awesome-claude-skills77k

webapp-testing

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

Browser automation

browser
code-yeongyu/oh-my-openagent70k

browser

Drives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch.

Browser automation

agent-browser
shanraisshan/claude-code-best-practice67k

agent-browser

Browser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking screenshots, extracting data, testing web apps, or automating any browser task. Triggers include requests to "open a website", "fill out a form", "click a button", "take a screenshot", "scrape data from a page", "test this web app", "login to a site", "automate browser actions", or any task requiring programmatic web interaction.

Browser automation

cherry-regression-test
CherryHQ/cherry-studio52k

cherry-regression-test

Run Cherry Studio critical-path system regression tasks through the repository-owned Playwright E2E workflow. Use for full regression, release acceptance, development-branch system validation, or a named cherry-regression-test task on GitHub-hosted macOS and Windows runners.

Browser automation