跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

ab-test-readout

Analyse a finished A/B test and write the readout — the result, whether it's statistically and practically significant, what it means, and the ship/no-ship call. Use when asked to analyse experiment results, write an A/B test readout, interpret test data, or decide whether to ship a variant. Produces a clear verdict with the lift and confidence, segment cuts, the risks (peeking, novelty, sample), and a recommendation. Distinct from planning a test — this reads results.

测试1.4kskills/ab-test-readout/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/mohitagw15856/pm-claude-skills/ab-test-readout/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

A/B Test Readout Skill

The hard part of an experiment is the readout: not "B won" but "is this real, is it big enough to matter, and should we ship?" This skill turns results into an honest decision — and flags the ways A/B results lie.

Working from a brief

Given results (even partial), write the full readout anyway. If significance isn't provided, reason about it from the numbers and flag what's needed to confirm. Mark assumed figures. Never declare a winner without addressing significance and sample.

Required Inputs

Ask for (if not already provided):

  • The hypothesis and the primary metric
  • Results — control vs variant: conversions/rate, sample size per arm, duration
  • Guardrail metrics (revenue, retention, latency, complaints) that mustn't regress
  • Pre-registered decision rule (what would count as a win) if one exists

Output Format

1. Verdict (one line)

Ship / Don't ship / Inconclusive — keep running — with the headline number.

2. The result

MetricControlVariantRelative liftSignificant?
Primaryp / CI
Guardrail(s)

State statistical significance (p-value / confidence interval) and practical significance (is the lift big enough to matter given the cost?).

3. Did it really win?

Address the ways A/B tests mislead:

  • Sample / power — was the test adequately powered, or under-sampled?
  • Peeking — was the call made early, inflating false positives?
  • Novelty / primacy — could the effect fade?
  • Segments — does the win hold across key segments, or is it driven by one?

4. Segment cuts

Where the effect is strong vs flat vs negative (new vs returning, platform, geography).

5. Recommendation & next step

Ship / iterate / re-run, plus what to monitor post-launch or what the follow-up test should isolate.

Quality Checks

  • Distinguishes statistical from practical significance
  • Checks guardrail metrics, not just the primary
  • Flags peeking, power, novelty, and segment-driven wins
  • Recommendation follows from the evidence, with a monitoring/next-test step
  • Doesn't declare a winner on an underpowered or peeked result

Anti-Patterns

  • "B won by 8%!" with no significance or sample size
  • Calling a result early (peeking) and shipping
  • Ignoring a guardrail regression because the primary went up
  • A statistically significant but practically meaningless lift treated as a win

Example Trigger Phrases

  • "Analyse experiment results."
  • "Write an A/B test readout."
  • "Interpret test data."
  • "Decide whether to ship a variant."

相似的 Skill

skill-creator
anthropics/skills180k

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

测试

ponytail-audit
DietrichGebert/ponytail158k

ponytail-audit

Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find bloat", "what can I delete", /ponytail-audit.

测试

ponytail-audit
DietrichGebert/ponytail158k

ponytail-audit

Quality audit of the whole repo: bugs, security, real load, missing tests, speed, and what to delete. Most important first.

测试

ponytail-review
DietrichGebert/ponytail158k

ponytail-review

Quality review of a diff: bugs, security, real load, missing tests, speed, and what to delete. Each finding says what goes wrong and how to fix it.

测试

ci-cd-and-automation
addyosmani/agent-skills103k

ci-cd-and-automation

Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.

测试

idea-refine
addyosmani/agent-skills103k

idea-refine

Refines raw ideas into sharp, actionable concepts through structured divergent and convergent thinking. Use when an idea is still vague, when you need to stress-test assumptions before committing to a plan, or when you want to expand options before converging on one. Triggers on "ideate", "refine this idea", or "stress-test my plan".

测试