跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

run-evals

do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals. Launches iPolloWork on Daytona or local Electron and runs the coded eval flows via CDP. Launch + run mechanics; the proof loop itself is the fraimz skill.

测试6.8k.opencode/skills/run-evals/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/devin-axis/ipollowork/run-evals/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Skill: Run Evals

Launch a real iPolloWork app and run coded eval flows against it. This skill owns launch + run; the prove/repair/verdict loop and evidence standard live in the fraimz skill — load that too for anything that ends in a verdict.

Prerequisites

  • daytona CLI installed and logged in (daytona login), right org selected (daytona organization use "<org-name>")
  • .devcontainer/ files present in the repo
  • Optional provider coverage: reusable secrets volume populated once with bash .devcontainer/setup-daytona-secrets-volume.sh .newtoken (never print keys; sandboxes source every /daytona-secrets/*.env before Electron starts)

Preferred path: Daytona sandbox

daytona organization use "<org-name>"
bash .devcontainer/test-on-daytona.sh <branch-or-commit> --artifacts-volume

The helper creates a fresh VNC-capable sandbox from the ipollowork-eval-vnc snapshot, mounts the secrets + pnpm-store volumes, starts XFCE/noVNC, Vite, and Electron with Daytona-safe flags, waits for CDP, then prints the CDP and noVNC URLs. --artifacts-volume mounts /daytona-artifacts served on port 8090 for published frame proof. Refresh the snapshot when dependencies change: bash .devcontainer/create-daytona-ipollowork-snapshot.sh.

Verify the endpoint before running flows:

curl -fsS "<CDP_URL>/json/list"   # must include an iPolloWork page target

If it fails, inspect /tmp/electron.log — the real success marker is Chromium's DevTools listening on ws://127.0.0.1:9825/....

If the app shows the Welcome page, use a coded onboarding flow or create /workspace/hello through the visible UI before running a workspace-dependent flow.

Run the flows

pnpm evals --list
pnpm evals --flow <flow-id> --cdp-url <printed-electron-cdp-url>
pnpm evals --all --stack den     # brings up MySQL + den-api + seed for cloud flows

The runner produces machine-checkable assertions, validated screenshots, and writes fraimz.html + report.md / report.json under evals/results/<run-id>/. If no coded flow exists for the behavior, add one in evals/flows/<id>.flow.mjs (see the fraimz skill and evals/README.md for the ctx.* API); use manual browser tools only to debug or prototype — a coded flow is the PR evidence.

Recording (motion only)

Frame proof is the default deliverable; record video only when motion matters (streaming, animations). Start with bash .devcontainer/test-on-daytona.sh <branch> --record-video --recording-name <name>, stop with daytona exec "$SANDBOX" -- 'bash .devcontainer/stop-daytona-recording.sh', download via the port-8090 artifacts URL. Details: daytona-recording-artifacts.

Local fallback

When Daytona is down or quota-limited:

pnpm install
pnpm --filter @ipollowork/app typecheck
IPOLLOWORK_ELECTRON_REMOTE_DEBUG_PORT=9826 pnpm dev   # then:
pnpm evals --flow <flow-id> --cdp-url http://127.0.0.1:9826

Report clearly whether the result came from Daytona or the local fallback — a local run is not a Daytona validation.

Teardown

daytona delete "$SANDBOX"

相似的 Skill

skill-creator
anthropics/skills180k

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

测试

ponytail-review
DietrichGebert/ponytail159k

ponytail-review

Quality review of a diff: bugs, security, real load, missing tests, speed, and what to delete. Each finding says what goes wrong and how to fix it.

测试

ponytail-audit
DietrichGebert/ponytail159k

ponytail-audit

Quality audit of the whole repo: bugs, security, real load, missing tests, speed, and what to delete. Most important first.

测试

ponytail-audit
DietrichGebert/ponytail159k

ponytail-audit

Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find bloat", "what can I delete", /ponytail-audit.

测试

doubt-driven-development
addyosmani/agent-skills103k

doubt-driven-development

Subjects every non-trivial decision to a fresh-context adversarial review before it stands. Use when you want every assumption cross-examined before proceeding, when stress-testing a plan for hidden failure modes, when correctness matters more than speed, when working in unfamiliar code, when stakes are high (production auth, security-sensitive logic, a high-stakes migration, irreversible operations), or any time a confident output would be cheaper to verify now than to debug later.

测试

idea-refine
addyosmani/agent-skills103k

idea-refine

Refines raw ideas into sharp, actionable concepts through structured divergent and convergent thinking. Use when an idea is still vague, when you need to stress-test assumptions before committing to a plan, or when you want to expand options before converging on one. Triggers on "ideate", "refine this idea", or "stress-test my plan".

测试