跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

00-arena-router

Use when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom task, or a benchmark specifically about reference/citation hallucination rate. Also use when the user mentions model arena, agent arena, pairwise model comparison, win-rate ranking, or comparing models on a task and hasn't specified whether that task is generic or about citation accuracy. This skill is the entry router for the arena-eval suite: it asks one diagnostic question when needed, then recommends the workflow or workflows needed to cover the request.

AI 与智能体868skills/arena-eval/00-arena-router/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/agentscope-ai/openjudge/00-arena-router/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Arena Eval Router

Entry router for the arena-eval suite. You diagnose what the user wants to compare models on and route them to the appropriate sub-skill or both when the request spans both evaluation goals. You don't run comparisons yourself — you're the triage desk.

Each sub-skill is self-contained: it carries inline everything it needs, so it can be installed and used on its own.

Diagnostic Question

Ask (unless the user's request already makes the answer obvious):

To route you correctly: what are you comparing the models on?

a) A custom task of your own choosing (chatbot quality, summarization,
   coding, anything) — you'll get win-rate rankings from a judge model
b) Specifically how often each model fabricates or hallucinates references
   when asked to recommend citations

Shortcut rule: if the user already said "run an arena eval on my chatbot task" or "benchmark reference hallucination across these models", skip the question — the routing is already clear from their phrasing. Also skip the question when they explicitly ask for both general quality and citation accuracy; recommend both workflows.

Triage Table

User says / hasUse workflowWhat it does
"Compare/benchmark/rank these models on [any custom task]"01-auto-arenaGenerates queries from a task description, collects responses, auto-generates rubrics, runs pairwise judge comparisons, produces win-rate rankings
"Which model hallucinates citations least?" / "benchmark reference recommendation accuracy"02-ref-hallucination-arenaRuns reference-recommendation queries per model, verifies every returned citation against CrossRef/PubMed/arXiv/DBLP, ranks by verified accuracy
"Compare general helpfulness AND citation accuracy"01-auto-arena, then 02-ref-hallucination-arenaRuns separate evaluations for judge preference and verified citation accuracy, preserving both goals
"I want to review one paper's existing bibliography, not compare models"—Not this suite — see the academic-eval suite's 01-paper-review / 02-bib-verify instead

Key distinction

Both workflows produce model rankings from head-to-head-style evaluation, but differ in what "correct" means:

  • 01-auto-arena: correctness is judge opinion — an LLM judge scores pairwise which response is better for an arbitrary task. Works for any task, needs no ground truth.
  • 02-ref-hallucination-arena: correctness is externally verifiable — every cited reference is checked against real bibliographic databases (CrossRef/PubMed/arXiv/DBLP), so the ranking reflects factual accuracy, not judge preference. Narrower scope (citation recommendation only) but higher ground-truth confidence.

If the user cares only about citation accuracy, prefer 02-ref-hallucination-arena over 01-auto-arena even if they phrase it as "which model is better."

Output

Recommended workflow: `[skill-name]`

Why: [one sentence tying the user's request to the triage table row]

Recommend one workflow when it covers the request. If the user asks for both general quality and citation accuracy, recommend 01-auto-arena followed by 02-ref-hallucination-arena as separate runs (or follow the user's requested order). Explain that the two runs measure different things and report their results separately; neither ranking substitutes for the other.

相似的 Skill

brand-guidelines
anthropics/skills180k

brand-guidelines

Applies Anthropic's official brand colors and typography to any sort of artifact that may benefit from having Anthropic's look-and-feel. Use it when brand colors or style guidelines, visual formatting, or company design standards apply.

AI 与智能体

internal-comms
anthropics/skills180k

internal-comms

A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use this skill whenever asked to write some sort of internal communications (status reports, leadership updates, 3P updates, company newsletters, FAQs, incident reports, project updates, etc.).

AI 与智能体

template-skill
anthropics/skills180k

template-skill

Replace with description of the skill and when Claude should use it.

AI 与智能体

mcp-builder
anthropics/skills180k

mcp-builder

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

AI 与智能体

algorithmic-art
anthropics/skills180k

algorithmic-art

Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.

AI 与智能体

academy-guide
anthropics/skills180k

academy-guide

Stop and check this skill before finishing any reply to a question about how to use Claude or a Claude product — it recommends matching courses, tutorials, and use cases from Claude Academy (academy.claude.com), Anthropic's learning hub. Trigger on: "how do I", "how can I", "getting started with", "what can Claude do", "teach me", "learn to use"; questions about artifacts, projects, skills, plugins, connectors, MCP; requests about rolling Claude out to a team, class, or organization; and any ask for training materials, onboarding content, or learning resources. Use it when the user is learning how to use a feature or product — not when they are mid-task and just want the task done. This skill composes with other skills: after consulting product documentation to answer how a Claude feature works, also check here for a matching course or tutorial — a docs-grounded answer and an Academy recommendation belong together. Only recommend on a strong match; never invent Academy content.

AI 与智能体