跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

gentle-ai-bench

Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis. Author and verify gentle-ai bench journeys; go test ./bench never proves driven execution.

测试7.6kskills/gentle-ai-bench/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/gentleman-programming/gentle-ai/skills-gentle-ai-bench/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Activation Contract

Load when touching bench/ in gentle-ai, adding or changing a journey, changing a product semantic a journey might pin, or diagnosing a bench failure in CI's Unit Tests job.

Hard Rules

  • go test ./bench validates corpus declarations only. It does NOT execute journeys. The only driven proof is building the harness and the product binary and running the harness against it; a green go test ./bench claims nothing about execution.
  • Reproduce CI, do not guess invocations: read the Unit Tests step in .github/workflows/ci.yml and copy its exact build and gentle-ai-bench run --binary ... commands. Use --only <journey-id> to drive one journey.
  • Journey IDs are unique across every journeys_*.go file. The collision guard fails loudly naming both files; pick an unused ID by reading the corpus, never reuse a retired one.
  • Every journey declares Review: — reviewOptedIn (the runner enables receipt-driven development globally before the first step, uncounted, and fails the journey if the switch does not come on) or reviewUntouched (its subject IS the switch, or it has nothing to do with reviews). The declaration is mandatory; validateCorpus fails the run without it. Lifecycle journeys must not depend on the product default. Reviews default to ON; reviewUntouched does not imply OFF. Journeys requiring OFF must explicitly disable it, while default-mode journeys must assert ON/default with unset sources.
  • Every execute transition must carry a runnable command; the dead-execute guard fails the run otherwise.
  • When a ratified product semantic changes, grep the corpus for journeys pinning the OLD behavior before shipping. The corpus is a second test surface beyond unit tests; a journey asserting the defect keeps the defect green.
  • dead_end prints n/a unless the run actually measured one. Never fabricate a value to move the column.
  • A by_design exemption costs a shape from the closed vocabulary plus a verified quote of the product's own next-action text. If the quote no longer tells the operator what to do, it is a defect wearing an exemption.
  • Prefer a NEW journeys_*.go file when the shared ones are owned by open PRs; bump the core journey-count pin in the same change.

Execution Steps

  1. Read the corpus area you touch and the CI invocation before writing.
  2. Author or adapt the journey; update its title, step names, and comment to say WHY the expectation holds (cite the issue or ratified decision).
  3. Run go test ./... in bench/ for declarations, THEN the driven harness for execution; both results go in the PR body.
  4. On semantic changes, list the journeys you checked for stale pins.

Output Contract

PR evidence includes the driven-mode summary line (completed / unsupported / failed counts) from a locally built binary, not only go test output.

相似的 Skill

skill-creator
anthropics/skills180k

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

测试

ponytail-review
DietrichGebert/ponytail159k

ponytail-review

Quality review of a diff: bugs, security, real load, missing tests, speed, and what to delete. Each finding says what goes wrong and how to fix it.

测试

ponytail-audit
DietrichGebert/ponytail159k

ponytail-audit

Quality audit of the whole repo: bugs, security, real load, missing tests, speed, and what to delete. Most important first.

测试

ponytail-audit
DietrichGebert/ponytail159k

ponytail-audit

Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find bloat", "what can I delete", /ponytail-audit.

测试

doubt-driven-development
addyosmani/agent-skills103k

doubt-driven-development

Subjects every non-trivial decision to a fresh-context adversarial review before it stands. Use when you want every assumption cross-examined before proceeding, when stress-testing a plan for hidden failure modes, when correctness matters more than speed, when working in unfamiliar code, when stakes are high (production auth, security-sensitive logic, a high-stakes migration, irreversible operations), or any time a confident output would be cheaper to verify now than to debug later.

测试

idea-refine
addyosmani/agent-skills103k

idea-refine

Refines raw ideas into sharp, actionable concepts through structured divergent and convergent thinking. Use when an idea is still vague, when you need to stress-test assumptions before committing to a plan, or when you want to expand options before converging on one. Triggers on "ideate", "refine this idea", or "stress-test my plan".

测试