Skill3.5k repo starsupdated today
wiki-dedup
wiki-dedup identifies and merges duplicate wiki pages covering the same concepts under different names. Use it in audit mode first to report candidate duplicates based on title and metadata similarity scoring, then proceed to merge mode only after confirming which pages should be consolidated. This skill is write-heavy and potentially destructive, requiring careful review before execution.
Install in Claude Code
Copygit clone --depth 1 https://github.com/Ar9av/obsidian-wiki /tmp/wiki-dedup && cp -r /tmp/wiki-dedup/.skills/wiki-dedup ~/.claude/skills/wiki-dedupThen start a new Claude Code session; the skill loads automatically.
Definition
SKILL.md
# Wiki Dedup — Identity Resolution and Page-Level Deduplication
You are finding and merging wiki pages that cover the same concept under different names. This is a write-heavy, potentially destructive skill — page merges cannot be automatically undone. Work carefully and confirm before acting in merge mode.
**Follow the Retrieval Primitives table in `llm-wiki/SKILL.md`.** The candidate-detection pass uses only frontmatter and titles (cheap). Only open full page bodies for confirmed candidate pairs.
## Before You Start
**Writing profile:** Before drafting or rewriting natural-language Markdown, read and apply the `Writing Profile Resolution` section in `llm-wiki/SKILL.md`. Framework schema, provenance, safety, and operation-specific requirements take precedence.
`WRITING.md` preferences apply only to newly drafted or rewritten natural-language Markdown; preserve source content and structured records.
1. **Resolve config** — follow the Config Resolution Protocol in `llm-wiki/SKILL.md` (inline `@name` override → walk up CWD for `.env` → global config → prompt setup). This gives `OBSIDIAN_VAULT_PATH` and `OBSIDIAN_LINK_FORMAT`.
2. Read `index.md` to get the full page inventory with one-line descriptions and tags.
3. Read `log.md` briefly — if a dedup run just happened, note what was already merged.
## Modes
| Mode | Flag | Behavior |
|---|---|---|
| **Audit** | *(default)* | Report candidates only — no writes |
| **Merge** | `--merge` | Show each confirmed pair, ask for confirmation before merging |
| **Auto-merge** | `--auto` | Merge all high-confidence pairs (`score ≥ 0.90`) non-interactively |
If the user doesn't specify, run in **Audit** mode and present findings before asking whether to proceed.
## Step 1: Build the Page Registry
Glob all `.md` files in the vault (excluding `_archives/`, `_raw/`, `.obsidian/`, `index.md`, `log.md`, `hot.md`, `_insights.md`, and any file that contains `redirects_to:` in its frontmatter — those are already merged redirect stubs).
For each remaining page, extract from frontmatter:
- `node_id` — relative path from vault root, without `.md`
- `title` — frontmatter `title` field
- `aliases` — frontmatter `aliases` list (may be absent)
- `tags` — frontmatter `tags` list
- `category` — directory prefix
Build a lookup table: `node_id → {title, aliases, tags, category, summary}`.
## Step 2: Detect Candidate Pairs
For every pair of pages in the registry, compute a **similarity score** using these signals:
### 2a. Title similarity signals
| Signal | How to assess | Max contribution |
|---|---|---|
| **Token overlap** | Jaccard similarity of lowercased title word-tokens (split on spaces, hyphens, underscores, punctuation) | 0.65 |
| **Edit distance** | Normalized edit distance on lowercased titles: `1 - (edits / max(len_a, len_b))` | 0.40 |
| **Substring containment** | One title is a substring of the other (e.g. "RSC" ⊂ "React Server Components") | 0.50 |
| **Alias cross-match** | Page A's title appears in page B's `aliases`, or vice versa | 0.65 |
Composite title score = `min(max(token_overlap, edit_distance, substring), 0.65) + alias_cross_bonus`.
You don't need exact arithmetic — make a confident judgement about degree of similarity.
**Title extraction note:** Some pages use YAML block scalars (`title: >-` or `title: |`). When the `title:` value is `>-`, `>`, `|`, or `|-`, the actual title is on the next indented line — read it from there. Never compare the literal string `>-` as a title.
### 2b. Semantic signals (cheap pass)
| Signal | Points |
|---|---|
| Same `category` directory | +0.10 |
| Tag overlap ≥ 3 shared tags | +0.15 |
| Tag overlap ≥ 2 shared tags | +0.05 |
| Same first tag (dominant tag) | +0.05 |
### 2c. Threshold
Flag pairs with composite score ≥ **0.75** as **candidates**. Pairs scoring 0.90+ are **high-confidence**.
Score ranges → confidence labels:
| Score | Label |
|---|---|
| ≥ 0.90 | HIGH — almost certainly the same concept |
| 0.75–0.89 | MEDIUM — likely the same, verify |
| 0.60–0.74 | LOW — possible abbreviation or specialisation; skip unless user asks |
Only carry HIGH and MEDIUM candidates into Step 3.
### 2d. Quick exit rule
If the vault has fewer than 10 pages, skip the pair loop and report "vault too small to have meaningful duplicates". If the vault has more than 500 pages, process candidates in batches of 50 pairs — pause and report progress between batches.
## Step 3: Semantic Verdict
For each candidate pair (sorted by score descending):
1. Read both pages in full (full page read — justified because candidate pool is small).
2. Ask: are these pages covering the **same concept**, or are they distinct?
Assign one of three verdicts:
| Verdict | Meaning |
|---|---|
| `merge` | Same concept — different name, abbreviation, alias, or accidental duplicate. Safe to merge. |
| `keep-separate` | Related but distinct — e.g. "Server Actions" vs "Server Components" are related React features, not duplicates. |
| `needs-review` | Ambiguous — substantial overlap but also meaningful differences. Flag for the user to decide. |
Attach a short reason to each verdict (one sentence). This appears in the report and the log.
## Step 4: Audit Report
Always produce this report, even in merge/auto-merge mode (so the user sees what will happen):
```markdown
## Wiki Dedup Report
### High-Confidence Candidates (score ≥ 0.90): N pairs
| Score | Page A | Page B | Verdict | Reason |
|---|---|---|---|---|
| 0.95 | `concepts/rsc.md` | `concepts/react-server-components.md` | merge | "RSC" is the abbreviation; both pages cover identical material |
| 0.91 | `entities/vaswani-2017.md` | `references/attention-is-all-you-need.md` | keep-separate | One is a person stub, one is a paper reference |
### Medium-Confidence Candidates (score 0.75–0.89): N pairs
| Score | Page A | Page B | Verdict | Reason |
|---|---|---|---|---|
| 0.82 | `concepts/fine-tuning.md` | `concepts/finetuning.md` | merge | Same concept, hyphenation variant |