Skip to content
FunCoding

Search

Search docs, Skills and MCP

wiki-dedup

Detect and merge wiki pages representing the same concept under different names. Use for page-identity deduplication and consolidation. Destructive merges require confirmation; not for link repair or general linting.

代码质量与审查3.5k.skills/wiki-dedup/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/ar9av/obsidian-wiki/wiki-dedup/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Wiki Dedup — Identity Resolution and Page-Level Deduplication

You are finding and merging wiki pages that cover the same concept under different names. This is a write-heavy, potentially destructive skill — page merges cannot be automatically undone. Work carefully and confirm before acting in merge mode.

Follow the Retrieval Primitives table in llm-wiki/SKILL.md. The candidate-detection pass uses only frontmatter and titles (cheap). Only open full page bodies for confirmed candidate pairs.

Before You Start

Writing profile: Before drafting or rewriting natural-language Markdown, read and apply the Writing Profile Resolution section in llm-wiki/SKILL.md. Framework schema, provenance, safety, and operation-specific requirements take precedence. WRITING.md preferences apply only to newly drafted or rewritten natural-language Markdown; preserve source content and structured records.

  1. Resolve config — follow the Config Resolution Protocol in llm-wiki/SKILL.md (inline @name override → walk up CWD for .env → global config → prompt setup). This gives OBSIDIAN_VAULT_PATH and OBSIDIAN_LINK_FORMAT.
  2. Read index.md to get the full page inventory with one-line descriptions and tags.
  3. Read log.md briefly — if a dedup run just happened, note what was already merged.

Modes

ModeFlagBehavior
Audit(default)Report candidates only — no writes
Merge--mergeShow each confirmed pair, ask for confirmation before merging
Auto-merge--autoMerge all high-confidence pairs (score ≥ 0.90) non-interactively

If the user doesn't specify, run in Audit mode and present findings before asking whether to proceed.

Step 1: Build the Page Registry

Glob all .md files in the vault (excluding _archives/, _raw/, .obsidian/, index.md, log.md, hot.md, _insights.md, and any file that contains redirects_to: in its frontmatter — those are already merged redirect stubs).

For each remaining page, extract from frontmatter:

  • node_id — relative path from vault root, without .md
  • title — frontmatter title field
  • aliases — frontmatter aliases list (may be absent)
  • tags — frontmatter tags list
  • category — directory prefix

Build a lookup table: node_id → {title, aliases, tags, category, summary}.

Step 2: Detect Candidate Pairs

For every pair of pages in the registry, compute a similarity score using these signals:

2a. Title similarity signals

SignalHow to assessMax contribution
Token overlapJaccard similarity of lowercased title word-tokens (split on spaces, hyphens, underscores, punctuation)0.65
Edit distanceNormalized edit distance on lowercased titles: 1 - (edits / max(len_a, len_b))0.40
Substring containmentOne title is a substring of the other (e.g. "RSC" ⊂ "React Server Components")0.50
Alias cross-matchPage A's title appears in page B's aliases, or vice versa0.65

Composite title score = min(max(token_overlap, edit_distance, substring), 0.65) + alias_cross_bonus.

You don't need exact arithmetic — make a confident judgement about degree of similarity.

Title extraction note: Some pages use YAML block scalars (title: >- or title: |). When the title: value is >-, >, |, or |-, the actual title is on the next indented line — read it from there. Never compare the literal string >- as a title.

2b. Semantic signals (cheap pass)

SignalPoints
Same category directory+0.10
Tag overlap ≥ 3 shared tags+0.15
Tag overlap ≥ 2 shared tags+0.05
Same first tag (dominant tag)+0.05

2c. Threshold

Flag pairs with composite score ≥ 0.75 as candidates. Pairs scoring 0.90+ are high-confidence.

Score ranges → confidence labels:

ScoreLabel
≥ 0.90HIGH — almost certainly the same concept
0.75–0.89MEDIUM — likely the same, verify
0.60–0.74LOW — possible abbreviation or specialisation; skip unless user asks

Only carry HIGH and MEDIUM candidates into Step 3.

2d. Quick exit rule

If the vault has fewer than 10 pages, skip the pair loop and report "vault too small to have meaningful duplicates". If the vault has more than 500 pages, process candidates in batches of 50 pairs — pause and report progress between batches.

Step 3: Semantic Verdict

For each candidate pair (sorted by score descending):

  1. Read both pages in full (full page read — justified because candidate pool is small).
  2. Ask: are these pages covering the same concept, or are they distinct?

Assign one of three verdicts:

VerdictMeaning
mergeSame concept — different name, abbreviation, alias, or accidental duplicate. Safe to merge.
keep-separateRelated but distinct — e.g. "Server Actions" vs "Server Components" are related React features, not duplicates.
needs-reviewAmbiguous — substantial overlap but also meaningful differences. Flag for the user to decide.

Attach a short reason to each verdict (one sentence). This appears in the report and the log.

Step 4: Audit Report

Always produce this report, even in merge/auto-merge mode (so the user sees what will happen):

## Wiki Dedup Report

### High-Confidence Candidates (score ≥ 0.90): N pairs

| Score | Page A | Page B | Verdict | Reason |
|---|---|---|---|---|
| 0.95 | `concepts/rsc.md` | `concepts/react-server-components.md` | merge | "RSC" is the abbreviation; both pages cover identical material |
| 0.91 | `entities/vaswani-2017.md` | `references/attention-is-all-you-need.md` | keep-separate | One is a person stub, one is a paper reference |

### Medium-Confidence Candidates (score 0.75–0.89): N pairs

| Score | Page A | Page B | Verdict | Reason |
|---|---|---|---|---|
| 0.82 | `concepts/fine-tuning.md` | `concepts/finetuning.md` | merge | Same concept, hyphenation variant |

### Needs Human Review: N pairs

| Score | Page A | Page B | Reason |
|---|---|---|---|
| 0.78 | `concepts/agents.md` | `concepts/autonomous-agents.md` | Substantial overlap but "agents" may intentionally be broader |

### Summary
- Pages scanned: N
- Candidate pairs found: M
- Recommended merges: X
- Keep separate: Y
- Needs review: Z

In Audit mode, stop here and ask: "Run --merge to interactively merge the recommended pairs, or --auto to merge all high-confidence ones automatically?"

Step 5: Merge

Pre-write snapshot — before the first file write, check whether the vault itself is the root of a Git repository. Merely being a subdirectory of a larger repository does not qualify: running git add -A there could capture unrelated files. If the vault is not a standalone Git repository, skip this step silently — no nagging, no suggesting git init.

VAULT_REAL_PATH=$(cd "$OBSIDIAN_VAULT_PATH" && pwd -P)
VAULT_GIT_ROOT=$(git -C "$OBSIDIAN_VAULT_PATH" rev-parse --show-toplevel 2>/dev/null || true)
SNAPSHOT_SHA=""

if [ -n "$VAULT_GIT_ROOT" ] && [ "$VAULT_GIT_ROOT" = "$VAULT_REAL_PATH" ]; then
  if git -C "$OBSIDIAN_VAULT_PATH" diff --quiet \
    && git -C "$OBSIDIAN_VAULT_PATH" diff --cached --quiet \
    && [ -z "$(git -C "$OBSIDIAN_VAULT_PATH" ls-files --others --exclude-standard)" ]; then
    SNAPSHOT_SHA=$(git -C "$OBSIDIAN_VAULT_PATH" rev-parse HEAD)
  else
    if ! git -C "$OBSIDIAN_VAULT_PATH" add -A; then
      echo "Pre-write snapshot failed; abort the skill without writing any vault files." >&2
      exit 1
    fi
    if ! git -C "$OBSIDIAN_VAULT_PATH" commit -m "pre-wiki-dedup snapshot" --quiet; then
      echo "Pre-write snapshot failed; abort the skill without writing any vault files." >&2
      exit 1
    fi
    SNAPSHOT_SHA=$(git -C "$OBSIDIAN_VAULT_PATH" rev-parse HEAD)
  fi
fi

The clean-repository branch deliberately avoids calling git commit, so "nothing to commit" is not treated as an error. If git add or git commit fails, stop before editing the vault; never continue without the promised snapshot.

If SNAPSHOT_SHA is non-empty and the skill writes files, include the SHA in the final report. To discard the entire run, after confirming there are no later changes worth keeping, the user can run:

git -C "$OBSIDIAN_VAULT_PATH" reset --hard "$SNAPSHOT_SHA"
git -C "$OBSIDIAN_VAULT_PATH" clean -fd

For each merge verdict pair (in merge or auto-merge mode):

In merge mode: show the pair and verdict, then ask: "Merge [Page A] into [Page B]? (yes/skip/review)". Skip on anything other than yes.

In auto-merge mode: only process HIGH-confidence (score ≥ 0.90) merges without prompting.

5a: Pick the canonical page

Apply these tiebreakers in order until one wins:

  1. More incoming wikilinks — grep the vault for [[node_id]] references; higher count wins
  2. Richer content — longer page body (more lines) wins
  3. More sources — larger sources: list wins
  4. Title length — longer, more descriptive title wins (e.g. "React Server Components" beats "RSC")
  5. Alphabetical — earlier title wins

The canonical page is the survivor. The other page becomes the secondary (to be merged in, then replaced with a redirect stub).

5b: Merge content into the canonical page

Read both pages. Update the canonical page:

  • aliases: — add secondary page's title and all its aliases (no duplicates)
  • tags: — merge both tag lists (deduplicate, cap at 5 domain tags + system tags)
  • sources: — merge both source lists (deduplicate)
  • relationships: — merge both relationship lists (deduplicate by target, prefer typed entries over untyped)
  • base_confidence — recompute using the union of sources and the formula from llm-wiki/SKILL.md
  • updated — set to now
  • summary: — rewrite to cover the merged scope if the secondary page added new ground
  • Body content — merge unique sections and bullets from the secondary page. Do not blindly append — integrate the content. Avoid duplicating claims already present in the canonical page. Use ^[inferred] markers where synthesis is needed.
  • provenance: — recompute after merging

5c: Write a redirect stub at the secondary page path

---
title: <secondary page title>
redirects_to: "[[<canonical node_id>]]"
aliases: [<secondary aliases>]
category: <secondary category>
tags: []
created: <secondary original created>
updated: <ISO timestamp now>
---

This page has been merged into [[<canonical page title>]].

The redirects_to: field tells any skill reading this page to follow the redirect rather than treat it as content.

Grep the entire vault for any link pointing at the secondary slug:

  • [[secondary-slug]] → [[canonical-slug]]
  • [[secondary-slug|display text]] → [[canonical-slug|display text]]
  • If OBSIDIAN_LINK_FORMAT=markdown: [text](../path/to/secondary.md) → [text](../path/to/canonical.md)

Safety rules:

  • Never rewrite inside code blocks (``` fences or inline code)
  • Never rewrite inside the redirect stub itself (that's the one place the old slug should remain legible)
  • Never use rm or destructive shell ops — only Edit/Write tools
  • Rewrite one file at a time, verifying each before moving on
  • If a file has zero occurrences, skip it

5e: Update tracking files

index.md and hot.md are regenerated by the memory sync call in Step 6 — the redirect stub is skipped by the page walker, so the secondary's entry disappears and the canonical entry picks up the merged summary on its own.

.manifest.json — For the secondary page's source entries: add "merged_into": "<canonical node_id>" to each. For the canonical page: merge in the secondary's pages_created and pages_updated lists.

5f: Final check

After all merges, grep the vault for any remaining [[secondary-slug]] references (in non-stub files). If any survive, report them — the rewrite step may have missed a non-standard link format.

Step 6: Log

One locked call writes the log line and reconciles the index and hot cache:

obsidian-wiki memory sync DEDUP \
  mode=<audit|merge|auto-merge> pages_scanned=<N> pairs_found=<M> \
  merged=<X> kept_separate=<Y> needs_review=<Z> \
  wikilinks_rewritten=<W> \
  --takeaways "Merged N duplicate pairs; canonical pages updated."

In audit mode (nothing merged) use obsidian-wiki memory log DEDUP ... instead — a read-only run must not rewrite the index or hot cache.

Never hand-edit index.md, log.md, or hot.md — the command takes the lock that keeps a parallel writer from dropping your update.

See .skills/llm-wiki/references/MEMORY.md for the full procedure.

Redirect Stub Handling

Other skills should handle redirect stubs as follows:

  • wiki-export — skip pages with redirects_to: in frontmatter; they are not content nodes
  • wiki-query — if a search hits a redirect stub, follow redirects_to: and read the canonical page instead
  • wiki-lint — validate that every redirects_to: wikilink resolves to an existing, non-stub page (a redirect chain — stub pointing to stub — is an error)
  • cross-linker — treat redirect stubs as non-targets; never add a new [[wikilink]] pointing at a stub page

Tips

  • Audit first, always. Even in auto-merge mode, the audit report is shown. Read it before trusting the results.
  • Check needs-review last. These are the hard cases — don't batch them with obvious merges.
  • Abbreviations are the most common case. "GPT" / "GPT-4" / "GPT4", "RSC" / "React Server Components", "LLM" / "Large Language Models" — these score high on substring containment and are almost always safe to merge.
  • Different versions are not duplicates. "GPT-3" and "GPT-4" are related but distinct. "fine-tuning" and "fine-tuning-llms" may be distinct (technique vs. specific application).
  • Run cross-linker after dedup. The redirect stubs leave the graph in a slightly inconsistent state. Cross-linker will tighten it up.

QMD Refresh After Vault Writes

QMD is a search index, not the source of truth. If $QMD_WIKI_COLLECTION is empty or unset, skip this step. Run it only after this skill has written or rewritten vault markdown. If QMD refresh fails, do not roll back the vault changes; report the QMD status separately.

Use $QMD_CLI if set; otherwise use qmd.

${QMD_CLI:-qmd} update

If the output says vectors are needed or embeddings may be stale, run:

${QMD_CLI:-qmd} embed

Verify the collection with either:

${QMD_CLI:-qmd} ls "$QMD_WIKI_COLLECTION"

or, when a specific page path is known:

${QMD_CLI:-qmd} get "qmd://$QMD_WIKI_COLLECTION/<page>.md" -l 5

Record one of:

  • QMD refreshed: update + embed + verified
  • QMD refreshed: update only + verified
  • QMD skipped: QMD_WIKI_COLLECTION unset
  • QMD skipped: qmd CLI unavailable
  • QMD failed: <short error summary>

Similar Skills

claude-api
anthropics/skills180k

claude-api

Reference for the Claude API / Anthropic SDK — model ids, pricing, params, streaming, tool use, MCP, agents, caching, token counting, model migration. TRIGGER — read BEFORE opening the target file; don't skip because it "looks like a one-liner" — whenever: the prompt names Claude/Anthropic in any form (Claude, Anthropic, Fable, Opus, Sonnet, Haiku, `anthropic`, `@anthropic-ai`, `claude-*`, `us.anthropic.*`, `[1m]`); the user asks about an LLM (pricing/model choice/limits/caching) — never answer from memory; OR the task is LLM-shaped with provider unstated (agent/MCP/tool-definition/multi-agent/RAG/LLM-judge/computer-use; generate/summarize/extract/classify/rewrite/converse over NL; debugging refusals/cutoffs/streaming/tool-calls/tokens). SKIP only when another provider is being worked on (overrides all triggers): OpenAI/GPT/Gemini/Llama/Mistral/Cohere/Ollama named in the query; OR `grep -rE 'openai|langchain_openai|google.generativeai|genai|mistralai|cohere|ollama'` over the project hits (run this grep FIRST if no provider named — don't Read the file).

Code quality & review

ponytail-review
DietrichGebert/ponytail158k

ponytail-review

Quality review of a change: is the logic right, is it safe, does it hold under real load, is risky code tested, is it fast enough, and is every line needed. Reads the connected code, not only the diff. Each finding is explained in plain English. Use for "review this", "code review", "review the last commit", "review my PR", "is this over-engineered", /ponytail-review.

Code quality & review

code-review-and-quality
addyosmani/agent-skills103k

code-review-and-quality

Conducts multi-axis code review. Use before merging any change. Use when reviewing code written by yourself, another agent, or a human. Use when you need to assess code quality across multiple dimensions before it enters the main branch. Use when asked to review a diff or a pull request, even when the diff is pasted inline.

Code quality & review

documentation-and-adrs
addyosmani/agent-skills103k

documentation-and-adrs

Records decisions and documentation. Use when you need to document an architecture decision (ADR) or the reasoning behind a design choice, when changing public APIs, shipping features, or when you need to record context that future engineers and agents will need to understand the codebase.

Code quality & review

code-simplification
addyosmani/agent-skills103k

code-simplification

Simplifies code for clarity. Use when refactoring code for clarity without changing behavior. Use when code works but is harder to read, maintain, or extend than it should be. Use when reviewing code that has accumulated unnecessary complexity.

Code quality & review

understand
Egonex-AI/Understand-Anything86k

understand

Analyze a codebase to produce an interactive knowledge graph for understanding architecture, components, and relationships

Code quality & review