Skip to content
FunCoding

Search

Search docs, Skills and MCP

evals-bootstrap

Scaffold a first eval suite for an agent: mine real failures into cases, write behavioural checks over traces, and generate the runner. Use when the user wants evals, regression tests for an agent, a golden set, or asks how to know a prompt/model change did not break things. Do NOT use for a single task's done-check (goal-test) or for auditing context (context-audit).

测试1kplugins/agents-course/skills/evals-bootstrap/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/undefined-ui/second-brain-os/evals-bootstrap/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Bootstrap the eval suite

Theory: Two kinds of checks and the full walkthrough in Evals practice. An eval suite is the same test after every change. Behavioural checks read the steps of a trace; end-to-end checks read only the result. Start behavioural: they are deterministic, run in seconds, and diagnose instead of just scoring.

Workflow

  1. Find real failures. Ask where the agent's runs live (logs, transcripts, a traces directory). Read until you have up to twenty real failures — not imagined ones. If there are no logged runs yet, build the trace logging first (step 3) and seed the suite with the three failures the user can recall; a small honest suite beats a large invented one.
  2. One line per failure. For each: the input, and the one specific behaviour that should have happened and did not. Failures cluster into four to eight behaviours; name them.
  3. Ensure traces exist. Each run must be stored as traces/<id>.json — a list of events including tool calls. If the user's harness is Claude Code, the transcript already is the trace; wire up whatever copies or converts it. No trace, no behavioural checks.
  4. Write cases.yaml. One entry per failure:
- id: refund_1042
  input: "Refund order #1042, customer says it arrived broken"
  expect: looks up the order before replying; asks approval before refund
  check:
    - trace has get_order before send_reply
    - trace has approval_request before refund

The expect line is for humans; the check lines are the test. Keep the rule language tiny: trace has X, trace has X before Y, trace lacks X. 5. Generate check_traces.py. A small runner: load cases.yaml, parse each rule with a regex, walk the tool-call list, print one line per case, exit non-zero on any failure. Keep it dependency-light (pyyaml only) and fast — the whole suite should run in seconds, with no model calls. 6. Run it and hand over the flywheel. Show the pass/fail lines. Then leave the loop in writing at the end of your report: read fresh traces weekly, add every new failure as a case, fix the biggest cluster, re-run.

Rules

  • Every case comes from a real failure; delete a case only when the behaviour it guards is retired, not when it is inconvenient.
  • No LLM-as-judge in the bootstrap. Add a judge later, only for what assertions cannot reach, and calibrate it against human labels first.
  • The suite must be one command (python check_traces.py) so it can gate a CI job or a pre-release habit without ceremony.

Similar Skills

skill-creator
anthropics/skills180k

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

Testing

ponytail-audit
DietrichGebert/ponytail158k

ponytail-audit

Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find bloat", "what can I delete", /ponytail-audit.

Testing

ponytail-audit
DietrichGebert/ponytail158k

ponytail-audit

Quality audit of the whole repo: bugs, security, real load, missing tests, speed, and what to delete. Most important first.

Testing

ponytail-review
DietrichGebert/ponytail158k

ponytail-review

Quality review of a diff: bugs, security, real load, missing tests, speed, and what to delete. Each finding says what goes wrong and how to fix it.

Testing

ci-cd-and-automation
addyosmani/agent-skills103k

ci-cd-and-automation

Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.

Testing

idea-refine
addyosmani/agent-skills103k

idea-refine

Refines raw ideas into sharp, actionable concepts through structured divergent and convergent thinking. Use when an idea is still vague, when you need to stress-test assumptions before committing to a plan, or when you want to expand options before converging on one. Triggers on "ideate", "refine this idea", or "stress-test my plan".

Testing