Skip to content
FunCoding

Search

Search docs, Skills and MCP

agent-design-review

Review an LLM agent design and find where it will be unreliable, expensive, or unsafe. Use when asked to review an agent architecture, critique a multi-step/tool-using agent, debug an agent that loops or goes off-task, or harden an agent before launch. Produces a structured review — task fit, control flow, tools, memory/context, failure handling, cost, and safety — with prioritised findings and fixes.

代码质量与审查1.4kskills/agent-design-review/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/mohitagw15856/pm-claude-skills/agent-design-review/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Agent Design Review Skill

Most agents don't fail because the model is weak — they fail because the design lets them loop, call the wrong tool, lose the thread across steps, or burn tokens with no stopping rule. This skill reviews an agent's architecture against the decisions that actually determine reliability, and ranks the fixes — so "it works in the demo but not in prod" becomes a specific list of changes. (Writing a new agent spec? Use agent-spec.)

Working from a brief

Given a sketch ("a research agent that searches, reads, and writes a report"), deliver the full review anyway — infer the likely control flow and tools, label the inference, and flag what to confirm. Never withhold the review for missing detail.

Required Inputs

Ask for these only if they aren't already provided (else infer and label):

  • What the agent does — its goal, and what a successful run produces.
  • Control flow — single prompt, plan-then-execute, ReAct loop, or multi-agent; and the stopping condition.
  • Tools & actions — what it can call, and which actions have side effects (write, send, pay).
  • Memory & context — what state carries across steps, and how context is kept in budget.
  • Constraints — latency, cost per run, and the trust boundary (untrusted input? real-world actions?).

Output Format

Agent Review: [agent]

1. Summary — will this be reliable in production? The top 3 risks and the single change that helps most.

2. Findings by dimension — for each, what's sound and what's fragile:

DimensionFindingSeverityFix
Control flowno max-steps / no progress check → loopsHighstep budget + "am I making progress?" check + halt
Tool useoverlapping tools confuse selectionMedfewer, sharply-described tools; allowlist
Contextfull history re-sent each step → cost + driftHighsummarise/scope memory per step
Failure handlingone tool error aborts the runMedretry/backoff + graceful degradation
Safetyacts without confirmation on writesHighhuman/confirm gate on side-effecting actions

3. Reliability checklist — termination guarantee (it always stops), error recovery, idempotency of side-effecting actions, and determinism where it matters.

4. Cost & latency — where tokens/steps are spent and how to cut them (cheaper model for sub-steps, caching, fewer round-trips) without losing quality. Pair with llm-cost-latency-budget.

5. Safety — untrusted input/tool output handled as data not instructions, least-privilege tools, and confirmation gates on high-impact actions. Pair with llm-guardrails-spec.

6. Prioritised fix plan — ordered by impact-to-effort.

Quality Checks

  • The agent has a guaranteed stopping condition (step/budget cap + progress check) — no unbounded loops
  • Side-effecting actions are idempotent or gated by a confirmation
  • Tools are few and sharply described so selection is unambiguous; access is least-privilege
  • Context strategy keeps the window in budget across steps (no naive full-history resend)
  • Tool errors are recovered, not fatal — retry/backoff and graceful degradation
  • Findings are severity-ranked and the fix plan is ordered by impact

Anti-Patterns

  • Do not approve an agent with no termination guarantee — "it usually stops" is an outage waiting to happen
  • Do not let it take irreversible actions without a confirmation gate
  • Do not give it many overlapping tools — selection accuracy drops as the toolset grows
  • Do not resend the whole history every step — cost and drift both climb
  • Do not treat tool/retrieved output as trusted instructions — it's the injection surface

Based On

LLM agent design practice — bounded control flow, least-privilege tool use, context management, error recovery, and safety gating.

Example Trigger Phrases

  • "Review an agent architecture."
  • "Critique a multi-step/tool-using agent."
  • "Debug an agent that loops."
  • "Harden an agent before launch."

Similar Skills

claude-api
anthropics/skills180k

claude-api

Reference for the Claude API / Anthropic SDK — model ids, pricing, params, streaming, tool use, MCP, agents, caching, token counting, model migration. TRIGGER — read BEFORE opening the target file; don't skip because it "looks like a one-liner" — whenever: the prompt names Claude/Anthropic in any form (Claude, Anthropic, Fable, Opus, Sonnet, Haiku, `anthropic`, `@anthropic-ai`, `claude-*`, `us.anthropic.*`, `[1m]`); the user asks about an LLM (pricing/model choice/limits/caching) — never answer from memory; OR the task is LLM-shaped with provider unstated (agent/MCP/tool-definition/multi-agent/RAG/LLM-judge/computer-use; generate/summarize/extract/classify/rewrite/converse over NL; debugging refusals/cutoffs/streaming/tool-calls/tokens). SKIP only when another provider is being worked on (overrides all triggers): OpenAI/GPT/Gemini/Llama/Mistral/Cohere/Ollama named in the query; OR `grep -rE 'openai|langchain_openai|google.generativeai|genai|mistralai|cohere|ollama'` over the project hits (run this grep FIRST if no provider named — don't Read the file).

Code quality & review

ponytail-review
DietrichGebert/ponytail158k

ponytail-review

Quality review of a change: is the logic right, is it safe, does it hold under real load, is risky code tested, is it fast enough, and is every line needed. Reads the connected code, not only the diff. Each finding is explained in plain English. Use for "review this", "code review", "review the last commit", "review my PR", "is this over-engineered", /ponytail-review.

Code quality & review

code-review-and-quality
addyosmani/agent-skills103k

code-review-and-quality

Conducts multi-axis code review. Use before merging any change. Use when reviewing code written by yourself, another agent, or a human. Use when you need to assess code quality across multiple dimensions before it enters the main branch. Use when asked to review a diff or a pull request, even when the diff is pasted inline.

Code quality & review

documentation-and-adrs
addyosmani/agent-skills103k

documentation-and-adrs

Records decisions and documentation. Use when you need to document an architecture decision (ADR) or the reasoning behind a design choice, when changing public APIs, shipping features, or when you need to record context that future engineers and agents will need to understand the codebase.

Code quality & review

code-simplification
addyosmani/agent-skills103k

code-simplification

Simplifies code for clarity. Use when refactoring code for clarity without changing behavior. Use when code works but is harder to read, maintain, or extend than it should be. Use when reviewing code that has accumulated unnecessary complexity.

Code quality & review

understand
Egonex-AI/Understand-Anything86k

understand

Analyze a codebase to produce an interactive knowledge graph for understanding architecture, components, and relationships

Code quality & review