Skip to content
FunCoding

Search

Search docs, Skills and MCP

senpi-qa

QA the omo Senpi adapter (packages/omo-senpi, packages/senpi-task) against the REAL senpi binary in strict isolation, and write every artifact to the one canonical evidence path .omo/evidence/omo-senpi-adapter/<slug>/. The live drivers under packages/omo-senpi/scripts/qa/ create their own isolated SENPI_CODING_AGENT_DIR and ignore the caller's, so the real ~/.senpi/agent is never written. Ships scripts/resolve-evidence-dir.mjs, which is the ONLY sanctioned way to pick an evidence directory: it rejects traversal, separators, absolute paths, and stray roots such as local-ignore/qa-evidence. Use whenever someone changes anything under packages/omo-senpi or packages/senpi-task, or wants to QA, smoke-test, verify, or debug the Senpi adapter, the task/team engine, the DAG, task RPC, or skill delivery. Triggers: senpi qa, qa senpi, senpi-qa, test senpi adapter, verify senpi task, senpi task e2e, senpi team e2e, task dag qa, live senpi driver, senpi evidence path.

测试70k.agents/skills/senpi-qa/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/code-yeongyu/oh-my-openagent/senpi-qa/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Senpi QA

QA the omo Senpi adapter (packages/omo-senpi/) and the task engine (packages/senpi-task/) by driving the REAL senpi binary. Unit tests never count as live QA here: bun run test:senpi is the package gate, the drivers in packages/omo-senpi/scripts/qa/ are the harness proof.

Golden rules

  • Evidence lives at exactly one path. Every artifact goes under .omo/evidence/omo-senpi-adapter/<slug>/. Pick it with scripts/resolve-evidence-dir.mjs and nothing else — a hand-typed path is how runs end up somewhere like local-ignore/qa-evidence/ or a .qa-evidence/ at the worktree root, which is outside the ignored root and gets committed by accident (#8703).
  • Evidence stays local. .omo/evidence/ is gitignored and the tracked-evidence audit test fails the build if any evidence path is tracked. Never git add -f an artifact; the PR body carries the summary and the decisive excerpts.
  • The real agent dir stays untouched. The live drivers build their own isolated SENPI_CODING_AGENT_DIR and deliberately IGNORE a caller-provided one, so ~/.senpi/agent is never used as the sandbox. Report the driver's realSenpiUntouched / changed-path fields and the isolated agent-dir path; treat a whole-directory digest as supporting evidence, not proof by itself.
  • No binary means SKIP, not silence. When senpi is absent the live drivers report SKIP or FAIL in their final JSON rather than degrading to the real home. A SKIP is not a pass — say so in the evidence README.
  • The captured JSON is the evidence. No file on disk means the QA did not happen, which means no commit and no push. The file proves the run on the machine that made it; it is not something the commit carries.

Resolve the evidence directory first

ev="$(node .agents/skills/senpi-qa/scripts/resolve-evidence-dir.mjs \
  --repo-root "$(git rev-parse --show-toplevel)" --slug <YYYYMMDD>-<short-slug>)"
mkdir -p "$ev"

The resolver returns an absolute path and creates nothing, so the caller decides when the directory appears. A slug is ONE relative segment of lowercase letters, digits, and hyphens (20260820-senpi-qa-contract). Separators, ./.., traversal, absolute paths, and a non-git root are rejected with a non-zero exit and a message naming the offending slug.

Router: pick your case

You changed…RunProves
Any adapter code, as the fast preconditionnode packages/omo-senpi/scripts/qa/drive.mjs --self-testthe driver + isolation harness itself works
Adapter wiring reaching a live sessionnode packages/omo-senpi/scripts/qa/drive.mjsa real senpi run with the plugin loaded, isolated agent dir, and no attributed real-home changes
Task lifecycle (single + batch)SENPI_BIN="$(command -v senpi)" node packages/omo-senpi/scripts/qa/task-e2e.mjslive task start/stream/terminal states
Team delivery, shutdown, reclaim, restart recoverySENPI_BIN="$(command -v senpi)" node packages/omo-senpi/scripts/qa/team-e2e.mjsinjection delivery and exactly-once recovery
Task RPC driver scriptsnode packages/omo-senpi/scripts/qa/task-rpc-e2e.mjs --self-testthe RPC surface contract
Skill delivery into a taskSENPI_BIN="$(command -v senpi)" node packages/omo-senpi/scripts/qa/task-load-skills-e2e.mjsskills reach the child
Continuation behaviornode packages/omo-senpi/scripts/qa/probe-continuation.mjsturns continue as expected
DAG state machine / runnersbun test packages/senpi-taskunit + chaos invariants (NOT live proof)

Point a driver's output at the resolved directory, e.g.:

TASK_E2E_OUT_DIR="$ev/live-task-dag" SENPI_BIN="$(command -v senpi)" \
  node packages/omo-senpi/scripts/qa/task-e2e.mjs

Package gate

tsgo --noEmit -p packages/omo-senpi/tsconfig.json
bun run test:senpi

Write the evidence README

Every run leaves $ev/README.md a reviewer can read without rerunning anything. The required sections are the repo-wide evidence rules in the root AGENTS.md (what was tested / observed / why it is enough / what was omitted). For Senpi, record the driver's changed-path/isolation fields and sandbox agent-dir path. Some drivers report sandbox paths without removing them; the caller must delete every task-owned sandbox and verify child PIDs are terminal before writing the cleanup receipt.

Similar Skills

skill-creator
anthropics/skills180k

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

Testing

ponytail-audit
DietrichGebert/ponytail158k

ponytail-audit

Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find bloat", "what can I delete", /ponytail-audit.

Testing

ponytail-audit
DietrichGebert/ponytail158k

ponytail-audit

Quality audit of the whole repo: bugs, security, real load, missing tests, speed, and what to delete. Most important first.

Testing

ponytail-review
DietrichGebert/ponytail158k

ponytail-review

Quality review of a diff: bugs, security, real load, missing tests, speed, and what to delete. Each finding says what goes wrong and how to fix it.

Testing

ci-cd-and-automation
addyosmani/agent-skills103k

ci-cd-and-automation

Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.

Testing

idea-refine
addyosmani/agent-skills103k

idea-refine

Refines raw ideas into sharp, actionable concepts through structured divergent and convergent thinking. Use when an idea is still vague, when you need to stress-test assumptions before committing to a plan, or when you want to expand options before converging on one. Triggers on "ideate", "refine this idea", or "stress-test my plan".

Testing