跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

codex-ab

Run an A/B codex review experiment — holistic codex review vs 3 focused dimension passes (security, ecto, liveview) on the branch diff, classify findings, report a panel-value verdict. Use when the branch is fresh, before any codex review runs.

安全565.claude/skills/codex-ab/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/oliver-kriska/claude-elixir-phoenix/codex-ab/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Codex Panel A/B (contributor instrument — verdict decided)

Answer one question with evidence: do dimension-focused codex passes find real issues that one holistic codex exec review misses? Runs both on the same diff, then classifies every focused finding against the holistic pass.

DECIDED 2026-07-10 after 4 runs (2 fresh): panel KILLED. Fresh-only 1 real miss / 1 false positive plus one zero-value run at 4× cost — real misses did not outnumber FPs. Kept as contributor tooling (NOT distributed) for one possible retest: a UI-heavy diff with a single extra liveview-focused pass (2× cost). Scoreboard: .claude/research/2026-07-03-codex-review-integration.md §7.

Usage

/codex-ab            # A/B against main (~5 min, 4 codex runs)
/codex-ab develop    # explicit base branch

Iron Laws

  1. FRESH DIFF ONLY — ask the user to confirm this branch has NOT been codex-reviewed yet (cloud or /phx:codex-loop). A drained diff returns NO FINDINGS everywhere and proves nothing — wasted quota
  2. Verify every REAL MISS in the code before counting it — a focused finding only scores if the issue actually exists at that file:line
  3. Read ONLY the findings .md files — streams are diverted to .log files; never cat a log into context (10k+ lines each)
  4. Exactly 4 codex runs, never re-run dimensions — bounded quota
  5. Persist the verdict — an unrecorded experiment is wasted quota

Workflow

Step 1: Preflight

Run command -v codex — missing → STOP with install hint. Then:

  • git status --short dirty → warn (codex flags local dirt as findings)
  • Ask: "Has codex already reviewed this branch (PR review or codex-loop)?" If yes → STOP, explain the fresh-diff requirement (Iron Law 1)

Step 2: Run the A/B (background, ~5 min)

bash ${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh {base} \
  .claude/reviews/codex-ab-$(date +%Y-%m-%d-%H%M)

Use run_in_background — it runs 1 holistic codex exec review + 3 focused codex exec workers (security / ecto / liveview) in parallel, all streams redirected. Do other work or wait; never poll.

Step 3: Classify

Read the 4 findings files (holistic.md, security.md, ecto.md, liveview.md — small). For EACH focused finding:

ClassMeaningTest
DUPLICATEHolistic already found itSame file + same defect
REAL MISSGenuine issue holistic missedRead the code at file:line — defect confirmed (Iron Law 2)
FALSE POSITIVEManufactured, pre-existing, or wrongCode check fails, or issue exists on base branch too

Step 4: Verdict

Present:

## Codex Panel A/B — {branch} vs {base}
| dimension | findings | duplicate | real miss | false positive |
Holistic-only findings: {n}
Verdict this run: {REAL MISS count} real miss vs {FP count} false positive
Decision rule: build --codex-panel only if real misses outnumber false
positives across 2-3 fresh branches.

Write the verdict table to .claude/reviews/codex-ab-{date}/VERDICT.md. Suggest repeating on the next 1–2 fresh branches before deciding.

Integration

fresh branch → /codex-ab (YOU ARE HERE) → verdict logged
   ├─ real misses win across runs → build /phx:review --codex-panel
   └─ duplicates/FPs win → keep holistic /phx:codex-loop, drop panel idea
       └─ OUTCOME 2026-07-10: this branch won — panel dropped

References

  • ${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh — the 4-run harness
  • Related: /phx:codex-loop (holistic fix loop), /phx:review --codex

相似的 Skill

security-and-hardening
addyosmani/agent-skills103k

security-and-hardening

Hardens code against vulnerabilities. Use when auditing an input handler for vulnerabilities, when handling user input, authentication, data storage, or external integrations, or when checking a login flow is safe against the OWASP Top Ten. Use when building any feature that accepts untrusted data, manages user sessions, or interacts with third-party services. Use when auditing dependencies for known vulnerabilities, triaging package-manager audit findings, or assessing supply-chain risk in a new package. Use when personal data or privacy compliance (GDPR, CCPA) is involved.

安全

archify
tt-a1i/archify80k

archify

Create polished, validated architecture, workflow, sequence, data-flow, and lifecycle/state diagrams as explorable standalone HTML with inline SVG, dark/light themes, optional trace motion, and PNG/JPEG/WebP/SVG/WebM export. Accept plain-language requirements or pasted Mermaid flowchart, sequenceDiagram, and stateDiagram input; inspect repository evidence when the diagram must reflect real code. Use when the user asks to visualize system architecture, infrastructure, cloud/security/network topology, technical workflows, API call sequences, request lifecycles, data pipelines, ETL/ELT, data lineage, state machines, or to convert/beautify Mermaid. Also use for everyday subjects with steps, parts, relationships, or states: a leave or travel plan, an application or approval process, a back-and-forth such as renting, where money or documents go, or where an application or order stands. Not for numeric charts or dashboards.

安全

security-research
code-yeongyu/oh-my-openagent70k

security-research

Team Mode security research skill. Orchestrates 3 vulnerability hunters and 2 PoC engineers to audit a codebase in parallel, prove exploitability, classify root causes, and calibrate severity by actual exploitability. Use for security review, vulnerability research, exploitability audit, pre-release security check, threat model validation, and `/security-research`. Triggers: 'security-research', 'security research', 'security review', 'vulnerability audit', 'exploitability audit', '보안 리뷰', '취약점 감사'.

安全

007
sickn33/agentic-awesome-skills47k

007

Security audit, hardening, threat modeling (STRIDE/PASTA), Red/Blue Team, OWASP checks, code review, incident response, and infrastructure security for any project.

安全

open-code-review
alibaba/open-code-review45k

open-code-review

Performs AI-powered code review on Git changes using the `ocr` CLI from alibaba/open-code-review. Use when the user asks to review code, review a pull request, review staged/unstaged changes, review a commit, or compare branches for code quality issues. Produces line-level review comments and can automatically apply fixes when requested. With appropriate review rules, can detect various types of issues including bugs, security vulnerabilities, performance problems, and code quality concerns.

安全

open-code-review
alibaba/open-code-review45k

open-code-review

Performs AI-powered code review on Git changes using the `ocr` CLI from alibaba/open-code-review. Use when the user asks to review code, review a pull request, review staged/unstaged changes, review a commit, or compare branches for code quality issues. Produces line-level review comments and can automatically apply fixes when requested. With appropriate review rules, can detect various types of issues including bugs, security vulnerabilities, performance problems, and code quality concerns.

安全