跳到正文
FunCoding

搜索

搜索文档、智能体、博客、Skill 和 MCP

senpi-qa

QA the omo Senpi adapter (packages/omo-senpi, packages/senpi-task) against the REAL senpi binary in strict isolation, and write every artifact to the one canonical evidence path .omo/evidence/omo-senpi-adapter/<slug>/. The live drivers under packages/omo-senpi/scripts/qa/ create their own isolated SENPI_CODING_AGENT_DIR and ignore the caller's, so the real ~/.senpi/agent is never written. Ships scripts/resolve-evidence-dir.mjs, which is the ONLY sanctioned way to pick an evidence directory: it rejects traversal, separators, absolute paths, and stray roots such as local-ignore/qa-evidence. Use whenever someone changes anything under packages/omo-senpi or packages/senpi-task, or wants to QA, smoke-test, verify, or debug the Senpi adapter, the task/team engine, the DAG, task RPC, or skill delivery. Triggers: senpi qa, qa senpi, senpi-qa, test senpi adapter, verify senpi task, senpi task e2e, senpi team e2e, task dag qa, live senpi driver, senpi evidence path.

测试70k.agents/skills/senpi-qa/SKILL.md

安装

将以下指令发送给 Claude Code、Codex 或 Cursor,智能体会先检查内容的安全性,经你确认后再安装。

读取 https://funcoding.ai/skills/code-yeongyu/oh-my-openagent/senpi-qa/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Senpi QA

QA the omo Senpi adapter (packages/omo-senpi/) and the task engine (packages/senpi-task/) by driving the REAL senpi binary. Unit tests never count as live QA here: bun run test:senpi is the package gate, the drivers in packages/omo-senpi/scripts/qa/ are the harness proof.

Golden rules

  • Evidence lives at exactly one path. Every artifact goes under .omo/evidence/omo-senpi-adapter/<slug>/. Pick it with scripts/resolve-evidence-dir.mjs and nothing else — a hand-typed path is how runs end up somewhere like local-ignore/qa-evidence/ or a .qa-evidence/ at the worktree root, which is outside the ignored root and gets committed by accident (#8703).
  • Evidence stays local. .omo/evidence/ is gitignored and the tracked-evidence audit test fails the build if any evidence path is tracked. Never git add -f an artifact; the PR body carries the summary and the decisive excerpts.
  • The real agent dir stays untouched. The live drivers build their own isolated SENPI_CODING_AGENT_DIR and deliberately IGNORE a caller-provided one, so ~/.senpi/agent is never used as the sandbox. Report the driver's realSenpiUntouched / changed-path fields and the isolated agent-dir path; treat a whole-directory digest as supporting evidence, not proof by itself.
  • No binary means SKIP, not silence. When senpi is absent the live drivers report SKIP or FAIL in their final JSON rather than degrading to the real home. A SKIP is not a pass — say so in the evidence README.
  • The captured JSON is the evidence. No file on disk means the QA did not happen, which means no commit and no push. The file proves the run on the machine that made it; it is not something the commit carries.

Resolve the evidence directory first

ev="$(node .agents/skills/senpi-qa/scripts/resolve-evidence-dir.mjs \
  --repo-root "$(git rev-parse --show-toplevel)" --slug <YYYYMMDD>-<short-slug>)"
mkdir -p "$ev"

The resolver returns an absolute path and creates nothing, so the caller decides when the directory appears. A slug is ONE relative segment of lowercase letters, digits, and hyphens (20260820-senpi-qa-contract). Separators, ./.., traversal, absolute paths, and a non-git root are rejected with a non-zero exit and a message naming the offending slug.

Router: pick your case

You changed…RunProves
Any adapter code, as the fast preconditionnode packages/omo-senpi/scripts/qa/drive.mjs --self-testthe driver + isolation harness itself works
Adapter wiring reaching a live sessionnode packages/omo-senpi/scripts/qa/drive.mjsa real senpi run with the plugin loaded, isolated agent dir, and no attributed real-home changes
Task lifecycle (single + batch)SENPI_BIN="$(command -v senpi)" node packages/omo-senpi/scripts/qa/task-e2e.mjslive task start/stream/terminal states
Team delivery, shutdown, reclaim, restart recoverySENPI_BIN="$(command -v senpi)" node packages/omo-senpi/scripts/qa/team-e2e.mjsinjection delivery and exactly-once recovery
Task RPC driver scriptsnode packages/omo-senpi/scripts/qa/task-rpc-e2e.mjs --self-testthe RPC surface contract
Skill delivery into a taskSENPI_BIN="$(command -v senpi)" node packages/omo-senpi/scripts/qa/task-load-skills-e2e.mjsskills reach the child
Continuation behaviornode packages/omo-senpi/scripts/qa/probe-continuation.mjsturns continue as expected
DAG state machine / runnersbun test packages/senpi-taskunit + chaos invariants (NOT live proof)

Point a driver's output at the resolved directory, e.g.:

TASK_E2E_OUT_DIR="$ev/live-task-dag" SENPI_BIN="$(command -v senpi)" \
  node packages/omo-senpi/scripts/qa/task-e2e.mjs

Package gate

tsgo --noEmit -p packages/omo-senpi/tsconfig.json
bun run test:senpi

Write the evidence README

Every run leaves $ev/README.md a reviewer can read without rerunning anything. The required sections are the repo-wide evidence rules in the root AGENTS.md (what was tested / observed / why it is enough / what was omitted). For Senpi, record the driver's changed-path/isolation fields and sandbox agent-dir path. Some drivers report sandbox paths without removing them; the caller must delete every task-owned sandbox and verify child PIDs are terminal before writing the cleanup receipt.

相似的 Skill

skill-creator
官方
anthropics/skills180k

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

测试

test-driven-development
addyosmani/agent-skills102k

test-driven-development

Drives development with tests using the red-green-refactor loop. Use when implementing any logic, fixing any bug, or changing any behavior. Use when you need to prove that code works, when a bug report arrives, or when you're about to modify existing functionality.

测试

interview-me
addyosmani/agent-skills102k

interview-me

Extracts what the user actually wants instead of what they think they should want. Achieves this through one-question-at-a-time interview until ~95% confidence about the underlying intent. Use when an ask is underspecified ("build me X" without "for whom" or "why now"), when the user explicitly invokes ("interview me", "grill me", "are we sure?", "stress-test my thinking"), or when you catch yourself silently filling in ambiguous requirements before any plan, spec, or code exists.

测试

idea-refine
addyosmani/agent-skills102k

idea-refine

Refines raw ideas into sharp, actionable concepts through structured divergent and convergent thinking. Use when an idea is still vague, when you need to stress-test assumptions before committing to a plan, or when you want to expand options before converging on one. Triggers on "ideate", "refine this idea", or "stress-test my plan".

测试

doubt-driven-development
addyosmani/agent-skills102k

doubt-driven-development

Subjects every non-trivial decision to a fresh-context adversarial review before it stands. Use when you want every assumption cross-examined before proceeding, when stress-testing a plan for hidden failure modes, when correctness matters more than speed, when working in unfamiliar code, when stakes are high (production auth, security-sensitive logic, a high-stakes migration, irreversible operations), or any time a confident output would be cheaper to verify now than to debug later.

测试

debugging-and-error-recovery
addyosmani/agent-skills102k

debugging-and-error-recovery

Guides systematic root-cause debugging. Use when tests fail, builds break, something that worked yesterday broke, behavior doesn't match expectations, or you encounter any unexpected error. Use when you need to figure out what broke and why — a systematic approach to finding and fixing the root cause rather than guessing.

测试