Skip to content
FunCoding

Search

Search docs, Skills and MCP

agent-self-eval

Post-run self-evaluation system that scores agent output on correctness, clarity, actionability, and conciseness. Use after /team runs, skill executions, or when explicitly asked to evaluate output quality.

扩展智能体511skills/agent-self-eval/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/coco-research/coco/agent-self-eval/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

@agents/PROMPT-DEFENSE.md

Agent Self-Evaluation

Score your own output (or another agent's output) across four axes to identify quality gaps and feed improvements into the learning system.

When to Use

  • After completing a /team:* pipeline run
  • After generating a deliverable (PRD, architecture doc, code review)
  • When user asks "how did I do?" or "evaluate this output"
  • Automatically at end of /gsd-execute-phase for quality tracking

Evaluation Axes

AxisQuestionFailure Signals
CorrectnessIs the output factually accurate and technically sound?Wrong APIs, broken references, hallucinated facts, logic errors
ClarityIs the explanation understandable and well-structured?Confusing structure, undefined jargon, missing context, rambling
ActionabilityCan the user act on the output immediately?Vague suggestions, missing steps, no verification path
ConcisenessDid it use the minimum tokens needed?Redundancy, over-explanation, filler content, restating the question

Scoring Scale

5 — Exceptional: no reasonable improvement possible
4 — Good: minor nits only, no substantive gaps
3 — Adequate: meets request but has notable weakness on ≥1 axis
2 — Weak: clear gap affecting usability or correctness
1 — Poor: fundamentally misses request or contains significant errors

The Evidence Rule

Every score below 5 MUST cite specific evidence. A score of 3 cannot just say "could be better" — it must say exactly what is missing or wrong. "Show the gap, don't just name it."

Procedure

Step 1: Collect Raw Material

Gather:

  • Original user request
  • Final output/deliverable
  • Tool outputs verifying correctness (test results, exit codes, lint)
  • User feedback received during task (corrections, "try again")

Step 2: Score Each Axis Independently

Rate 1-5 with mandatory evidence for scores <5.

Step 3: Generate Eval Report

SELF-EVALUATION REPORT
======================
Task: {brief description}
Overall: {weighted average}/5

CORRECTNESS: {score}/5
  Evidence: {specific finding or "No issues found"}

CLARITY: {score}/5
  Evidence: {specific finding or "No issues found"}

ACTIONABILITY: {score}/5
  Evidence: {specific finding or "No issues found"}

CONCISENESS: {score}/5
  Evidence: {specific finding or "No issues found"}

IMPROVEMENT INSTINCTS:
- {trigger} → {action} (confidence: {0.3-0.9})

Step 4: Feed Learning System

If learning system is active (PR-27+), auto-generate instinct YAML from findings:

---
id: eval-{task-slug}-{axis-lowercase}
trigger: "when {task type}"
action: "{specific improvement}"
confidence: 0.6
domain: quality
source: self-eval
scope: project
---

Integration Points

  • /team:verify invokes self-eval on Layer 2 output before Layer 3 review
  • /gsd-execute-phase runs self-eval per subagent, aggregates in SUMMARY.md
  • High-confidence eval instincts (≥0.8) auto-update team-toolkit.md quality notes
  • Low scores (≤2) on correctness trigger automatic re-execution offer

Similar Skills

skill-creator
anthropics/skills180k

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

Extending agents

template-skill
anthropics/skills180k

template-skill

Replace with description of the skill and when Claude should use it.

Extending agents

mcp-builder
anthropics/skills180k

mcp-builder

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

Extending agents

claude-api
anthropics/skills180k

claude-api

Reference for the Claude API / Anthropic SDK — model ids, pricing, params, streaming, tool use, MCP, agents, caching, token counting, model migration. TRIGGER — read BEFORE opening the target file; don't skip because it "looks like a one-liner" — whenever: the prompt names Claude/Anthropic in any form (Claude, Anthropic, Fable, Opus, Sonnet, Haiku, `anthropic`, `@anthropic-ai`, `claude-*`, `us.anthropic.*`, `[1m]`); the user asks about an LLM (pricing/model choice/limits/caching) — never answer from memory; OR the task is LLM-shaped with provider unstated (agent/MCP/tool-definition/multi-agent/RAG/LLM-judge/computer-use; generate/summarize/extract/classify/rewrite/converse over NL; debugging refusals/cutoffs/streaming/tool-calls/tokens). SKIP only when another provider is being worked on (overrides all triggers): OpenAI/GPT/Gemini/Llama/Mistral/Cohere/Ollama named in the query; OR `grep -rE 'openai|langchain_openai|google.generativeai|genai|mistralai|cohere|ollama'` over the project hits (run this grep FIRST if no provider named — don't Read the file).

Extending agents

academy-guide
anthropics/skills180k

academy-guide

Stop and check this skill before finishing any reply to a question about how to use Claude or a Claude product — it recommends matching courses, tutorials, and use cases from Claude Academy (academy.claude.com), Anthropic's learning hub. Trigger on: "how do I", "how can I", "getting started with", "what can Claude do", "teach me", "learn to use"; questions about artifacts, projects, skills, plugins, connectors, MCP; requests about rolling Claude out to a team, class, or organization; and any ask for training materials, onboarding content, or learning resources. Use it when the user is learning how to use a feature or product — not when they are mid-task and just want the task done. This skill composes with other skills: after consulting product documentation to answer how a Claude feature works, also check here for a matching course or tutorial — a docs-grounded answer and an Academy recommendation belong together. Only recommend on a strong match; never invent Academy content.

Extending agents

ponytail-help
DietrichGebert/ponytail160k

ponytail-help

Quick reference for ponytail levels, skills and commands. One-shot display. Use for /ponytail-help, "ponytail help", "how do I use ponytail".

Extending agents