跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

knowledge-base-management

管理本地 Markdown/Obsidian 知识库:素材入库、健康检查、增量索引、关键词与 semantic-lite 检索、逐字原文资料包、归档和 MCP 连接。用于搜索知识库、整理资料及为 Agent 准备可定位的引用。

文档与办公1.2kknowledge-base-management/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/chubbyguan/chubbyskills/knowledge-base-management/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Knowledge Base Management — 知识库全生命周期管理

核心架构

三层架构

素材库/  → Layer 1: 不可变原始素材(AI 只读)
wiki/    → Layer 2: AI 编译的结构化知识(AI 维护)
产出/    → Layer 3: 按需生成的视图(不持久化)

核心原则:素材库是事实源(不可变),Wiki 是 AI 编译的投影(必须溯源),产出是按需生成的视图(不持久化)。

Vault 配置

vault 路径通过环境变量配置,不写死在脚本里:

export VAULT_DIR="$HOME/Documents/your-vault"   # 你的 Obsidian vault 根目录

完整仓库可以统一配置采集、搜索和资料包:

python3 tools/chubby.py init --vault "$VAULT_DIR"

这里的路径是知识库根目录。采集进入根目录的 00_Inbox,默认索引为 .chubby/index.sqlite。独立安装本 skill 后,tools/vault_index.py 和 tools/evidence_brief.py 直接使用 VAULT_DIR 或 --vault,不读取仓库中的 chubby.yaml。

健康检查可自行加入调度器;安装本 skill 不会创建 cron。


1. 素材入库 (Ingest)

新素材处理流水线

用户丢素材 →
  ① 判断类型 → 存入素材库/对应子目录
  ② 提取关键信息 → 决定是否建 wiki 页面
  ③ 如果值得:建选题(打 SHARP 分)或更新项目页
  ④ 更新 wiki/index.md 索引

素材分级 (ABC Grading)

等级标准处理
A核心方法论/框架/系统优先编译为选题
B有洞察但不够系统备选,作为 A 的补充
C水货/重复/与定位无关归档,不主动编译

公众号文章处理

来源一:可选的 wechat-article-exporter 外部同步 如果已自行部署该工具和调度器,可把公众号文章同步到 素材库/公众号文章/<公众号名>/。本 skill 不内置该服务或定时任务。

来源二:手机保存的文章导入 通过手机保存的文章先落到一个待处理目录(按日期),再由一个归类脚本搬入素材库并按领域筛选优质文章做 A+B 深度处理。

归类脚本(如 sync_notes_to_kb.py)与各人的目录结构强相关,本仓库不内置,按下面的「工作流」自建即可。

工作流:

新文章入库 →
  ① 脚本自动归类到 素材库/公众号文章/<公众号名>/
  ② URL 去重(跳过已存在的文章)
  ③ 领域关键词筛选(AI/半导体/消费/电商/品牌/投资/Agent/Skill)
     跳过纯营销、与用户领域无关的文章
  ④ 命中关键词的精选文章 → A+B 双轨处理(每日上限 5 篇)
     A层:观点提取+分类+3条Takeaway → wiki/📡 外部输入/公众号/<公众号名>/<主题>/
     B层:问题链+伴读引导 → 同上
  ⑤ 更新 wiki/index.md(外部输入 — 最新 A+B 处理 表格)

关键原则:

  • 聚焦高价值主题(匹配用户兴趣),不追求全量覆盖
  • 每个 wiki 页面必须用 source frontmatter 列出引用素材路径
  • 深度处理上限 5 篇/天,分批消化
  • 文章路径格式:[[../../../../素材库/公众号文章/<公众号>/<文件名>]]

2. 健康检查 (Audit)

自动检查项

手动运行脚本或配置自己的调度器后,检查:

  • 断链(broken wikilinks)
  • 缺 source 字段
  • frontmatter 缺失
  • 待沉淀概念
  • 低链接密度

脚本:scripts/vault_health_check.py(本仓库提供,零依赖,纯标准库)

# 指定 vault 路径
python3 scripts/vault_health_check.py "$HOME/Documents/your-vault"

# 或读 VAULT_DIR 环境变量
export VAULT_DIR="$HOME/Documents/your-vault"
python3 scripts/vault_health_check.py --json

# 把缺 source 字段也视为问题(强溯源场景)
python3 scripts/vault_health_check.py --require-source

检查项:断链、缺 frontmatter/source/summary、孤立笔记、空目录、内容重复、文件名含空格。 存在断链时退出码为 1,方便接入 cron / CI 告警。

清理标准

操作规则
去重SHA256 或 diff 确认相同后删除副本
归位散落文件按类型移到素材库/对应子目录
归档历史版本移到 wiki/🗄️ 归档/,不删除
空目录直接删除。含 redirect README 的 legacy 目录也算空目录
全库审计执行第 6 节「全库审计与批量归档」
目录审查执行第 5 节「目录审查与增量清理」

文件命名规范

✅ 选题-AI-Agent-未来方向.md    # 连字符连接,无空格
❌ 选题-DemisHassabis Agent.md  # 含空格,wikilink 断裂

迁移提示页格式

当文档内容迁移到新位置时,旧文件不删除,改为迁移提示:

---
title: ⚠️ 已迁移 - 原文件名(旧版 vX.0)
type: note
tags: [已废弃, 旧版, 已迁移]
created: YYYY-MM-DD
summary: 已被 [[新位置/新文件名|新版本]] 取代。
---

关键规则:

  • frontmatter: ⚠️ 已迁移 标题 + 已废弃 标签
  • 说明为什么旧版不再适用(1-3条)
  • 明确的新链接
  • 不要保留旧的正文内容

3. 工具集成 (Tools)

🧠 增量索引与统一搜索

完整仓库的统一入口:

python3 tools/chubby.py init --vault "$VAULT_DIR"
python3 tools/chubby.py ingest "替换为真实素材链接" --no-enrich
python3 tools/chubby.py search "AI Agent"
python3 tools/chubby.py search "内容策略" --mode lite

成功采集后,以及统一搜索前,自动同步索引。init --vault 指向根目录;旧的 ingest --vault 参数仍表示具体入库目录。

独立安装本 skill 后,在 skill 目录内使用自带工具:

export VAULT_DIR="$HOME/Documents/your-vault"
python3 tools/vault_index.py index "$VAULT_DIR"
python3 tools/vault_index.py search "AI Agent"
python3 tools/vault_index.py semantic "内容策略"
python3 tools/vault_index.py search "品牌" --platform wechat
python3 tools/vault_index.py recent --limit 10
python3 tools/vault_index.py read "10_Sources/x/example.md" --vault "$VAULT_DIR"
python3 tools/vault_index.py stats

独立 vault_index.py 搜索读取已有索引,笔记变化后先再运行 index。后续命令需要保持 VAULT_DIR,或显式传入同一个数据库:

python3 tools/vault_index.py --db /path/to/index.sqlite index "$VAULT_DIR"
python3 tools/vault_index.py --db /path/to/index.sqlite search "内容策略"

SQLite FTS5 可用时自动启用;不可用时使用 LIKE 搜索。默认 semantic-lite 只做本地词语和字符组合排序,不调用模型 API。

索引规则:

  • 默认位置为 $VAULT_DIR/.chubby/index.sqlite;VAULT_INDEX_DB 或显式 --db 可以覆盖。
  • index 默认增量同步。未变笔记的向量保留,实际 embedding 输入变化时使旧向量失效,删除笔记时移除对应索引和向量。
  • 新默认索引不存在时,旧 $VAULT_DIR/.chubby/vault_index.sqlite 会复制迁移,原文件保留。其他位置的旧数据库需显式指定并验证所属知识库。
  • 数据库绑定知识库,扫描或读取失败会回滚同步,越出知识库范围的符号链接会被拒绝。
  • 全量重建需显式执行 python3 tools/vault_index.py index "$VAULT_DIR" --rebuild,会丢弃已有 embedding。

📥 导入本地材料

完整仓库中,python3 tools/chubby.py import "/path/to/document.md" 将 Markdown、文本或文字层 PDF 接入现有采集、复用和索引流程。可用 --source-url 指定原始网页;本地附件会复制并改写引用,原文件保留。PDF 需要可选 pymupdf,扫描件不会被当作成功导入。

独立安装本 skill 后,也可在 skill 目录使用:

python3 tools/import_document.py "/path/to/document.md" --output "$VAULT_DIR/00_Inbox"
python3 tools/vault_index.py index "$VAULT_DIR"

独立导入工具不写统一 CLI 的运行记录;MCP 查询或 brief 生成前仍会自动同步知识库。安装包包含该工具及所需公共模块。

📎 逐字原文资料包

完整仓库使用统一入口:

python3 tools/chubby.py brief --topic "内容策略" --output reports/content-brief.md

独立安装后,在 skill 目录运行:

python3 tools/evidence_brief.py --vault "$VAULT_DIR" --topic "内容策略" \
  --output "$VAULT_DIR/30_Output/content-brief.md"

资料包工具先自动同步索引,关键词检索没有命中时回退到本地 semantic-lite;生成的 research_brief 文档从候选中排除。输出 Markdown 与同名 JSON,包含逐字摘录、原始 source、笔记相对路径、真实文件行号和 SHA-256。

行号从源文件第一行计算,包含 frontmatter。导出前再次检查文件和摘录,来源变动时需要重新生成。已有输出默认拒绝覆盖,显式 --force 才替换。没有匹配内容时写明证据不足。

工具只整理原文证据包和 Agent 任务说明,选题由 Agent 另行生成;文件和摘录核验不代表原文事实已证实。要求 Agent 区分原作者观点与推断,并在引用中保留证据编号、路径和行号。

🗂️ 自动归档与知识卡片

tools/vault_curator.py 默认预览,确认后用 --apply 执行:

python3 tools/vault_curator.py archive "$VAULT_DIR"
python3 tools/vault_curator.py archive "$VAULT_DIR" --apply
python3 tools/vault_curator.py card "$VAULT_DIR" "10_Sources/x/example.md" --apply
  • archive 只处理 00_Inbox/**/*.md,按平台、摘要和 processed 标签归入 10_Sources/<platform>/ 或 20_Processed/。
  • card 生成 20_Processed/Cards/*.md,保留来源、摘要、要点和 tags。

🔌 MCP Server

scripts/mcp_server.py 为支持 MCP 的 Agent 提供搜索、读取和索引工具。使用 Python 3.10 或更新版本,在完整仓库根目录执行:

python3 -m pip install -r knowledge-base-management/requirements-mcp.txt
python3 tools/mcp_smoke.py --json
VAULT_DIR=/path/to/your-vault python3 knowledge-base-management/scripts/mcp_server.py

当前验证 SDK 为 mcp==1.30.0。mcp_smoke.py 用临时 vault 启动真实 stdio 服务,验证握手、6 个工具发现、搜索和读取,不访问你的知识库。--help 不需要安装 SDK。

完整仓库的安装器会一起复制 vault_index.py、vault_curator.py、evidence_brief.py、import_document.py 及所需公共模块:

python3 tools/install_skill.py knowledge-base-management --dest ~/.codex/skills
python3 -m pip install -r ~/.codex/skills/knowledge-base-management/requirements-mcp.txt
VAULT_DIR=/path/to/your-vault python3 ~/.codex/skills/knowledge-base-management/scripts/mcp_server.py

目标技能目录必须不存在;已有安装先备份,或使用新的安装目录。MCP 优先加载 skill 内的工具,安装后不依赖原仓库位置。

客户端配置使用真实的 Python 和 server 绝对路径:

{
  "mcpServers": {
    "chubby-kb": {
      "command": "/absolute/path/to/venv/bin/python",
      "args": ["/absolute/path/to/knowledge-base-management/scripts/mcp_server.py"],
      "env": { "VAULT_DIR": "/path/to/your-vault" }
    }
  }
}

工具名保持 6 个:search_vault、semantic_search_vault、read_kb_note、list_recent_notes、reindex_vault、vault_index_stats。搜索、语义检索、最近笔记和统计前自动同步索引;read_kb_note 直接读取原文;reindex_vault() 名称不变,执行增量同步。

默认数据库与 CLI 一致,为根目录下 .chubby/index.sqlite;自定义时在客户端环境中设置 VAULT_INDEX_DB。读取路径必须位于 VAULT_DIR 范围内。


以下均为可选的第三方/外部工具,不随本仓库提供。路径(如 ~/graphrag-poc/、 ~/.hermes/...)是作者本机的示例位置,请替换为你自己的安装路径。没有它们也不影响 第 1、2 节的核心流程与 scripts/vault_health_check.py。

知识库三件套

工具定位适用场景
GBrain搜索知道要找什么关键词
GraphRAG发现不知道关键词,想发现隐藏关联
LLM Wiki写作用 Karpathy 模式建 LLM 可读的 Wiki
Understand-Anything可视化看懂全库结构,发现隐性关联

GBrain 搜索集成

日常使用:

alias gs='GBRAIN_SKIP_RECIPES=1 gbrain search'
gs "定价 品牌"
gs "奥德赛时期"

# 混合搜索
gbrain query "消费者心理剩余"

# 读全文(注意 slug 含 emoji 目录名)
gbrain get "wiki/📡 外部输入/播客/XXX/观点提取"

# 查看统计
gbrain stats

配置:DeepSeek downstream LLM(deepseek_api_key + deepseek_model: deepseek-chat + search.mode: balanced) 完整手册:wiki/🧠 AI系统/GBrain-完整操作手册.md

GraphRAG 发现

⚠️ 已移至独立项目:按你的实际安装路径配置 Flask API:localhost:8999(默认) 内容:自定义知识库 + DeepSeek + BGE 嵌入

用于发现知识库中的隐性关联和跨领域连接。

LLM Wiki(Karpathy 模式)

Karpathy 的 LLM Wiki 模式:构建 interlinked markdown 知识库,让 LLM 在推理时直接查。 已集成到 Obsidian vault。

Understand-Anything — 知识图谱可视化

已安装为独立 skill。

使用:

# 必须指向 wiki 根目录,必须用 python3.11+
python3 parse-knowledge-base.py /path/to/your-vault/wiki

⚠️ 关键 pitfalls:

  • Python 3.11+ 必需
  • 必须指向 wiki 根目录(子目录没有 index.md 会报错)
  • Dashboard 需先 build core 包
  • 需要 GRAPH_DIR 和 ACCESS_TOKEN 环境变量

4. 目录整理 (Organization)

目录审查与增量清理

当用户说"看一下目录是否合理"时执行:

Step 1: 全量目录扫描
  → find . -maxdepth 3 -type d | sort
  → 对每个 wiki 子目录统计 .md 文件数

Step 2: 逐项检查
  | 检查项 | 特征 | 严重性 |
  |--------|------|--------|
  | 旧版残留目录 | diff -rq 确认新旧版本内容不同 | 🔴 |
  | Legacy redirect 只剩 README | 目录只有一个 redirect 文件 | 🟡 |
  | 概念卡片散落根目录 | wiki/🧬 知识图谱/XXX.md 应在概念/下 | 🟡 |
  | cron 脚本写入旧路径 | grep -rn "旧目录" ~/.hermes/scripts/ | 🔴 |

Step 3: 分类执行 → 归档/删除/移动/更新引用

Step 4: 验证 → 零残留

全库审计与批量归档

Step 1: 全量扫描目录结构和文件数
Step 2: 逐目录与最新管理办法对照
Step 3: 分级标记(🔴明确无用 → 🟡重复 → 🟠草稿 → 🔵过时)
Step 4: 批量 mv 到 wiki/🗄️ 归档/
Step 5: 修复残留引用(README/wiki/index/项目索引)

归档目录结构:

wiki/🗄️ 归档/
├── 归档清单-YYYY-MM-DD.md
├── 旧版管理办法/
├── 旧版流程/
├── 旧版草稿/
├── 旧选题/
└── ...

反模式:

  • ❌ 不要删除归档目录下的文件
  • ❌ 归档后必须同步更新所有 index/README 引用
  • ❌ 素材库/ 下的文件不归档
  • ❌ 归档时不要改文件名
  • ✅ 每次大清理必须创建 归档清单-YYYY-MM-DD.md

5. Wiki 实体管理

Wiki 实体格式

每个团队/项目在 wiki/teams/ 下有对应 .md 实体:

---
title: 团队名
created: YYYY-MM-DD
updated: YYYY-MM-DD
type: entity
tags: [team, ...]
sources: [50.团队/实际目录名]   # 重要:路径必须是实际存在的目录名
---

# 团队名

## Overview
## 文件结构(表格,列出子目录和数量)
## Related

sources 路径必须与实际目录名完全匹配:

  • ❌ sources: [50.团队/02.事业线] (旧/错)
  • ✅ sources: [50.团队/事业线] (新/对)

异常文件处理

某些 .md 文件可能被误认为目录名(如 千金-日报索引.md、OKR看板.md)。检查方式:

# 检查是否是文件被当作目录
for d in team_dir.iterdir():
    if d.is_file() and not d.name.startswith("."):
        print(f"⚠️ {d.name} 是文件却像目录")

修复:用 mv 把它们移到对应目录下。


6. 批量修复模式

问题:播客 wiki 文件中的 wikilink 引用素材库 EP 文件,因文件名含特殊字符(引号、emoji、反斜杠)导致相对路径解析失败。

修复方式:

# 1. 建索引: EP编号 → 实际文件路径
ep_index = {}
for f in 素材库_dir.iterdir():
    ep_match = re.match(r'(EP\d+)', f.name)
    if ep_match:
        ep_index[ep_match.group(1)] = f

# 2. 扫描 wiki 文件,用 EP 编号匹配实际文件,重建相对路径
for link in wikilinks:
    ep_match = re.search(r'(EP\d+)', target)
    if ep_match and ep_match.group(1) in ep_index:
        rel = os.path.relpath(ep_index[ep_id], md.parent)
        # 替换 wikilink

批量补 frontmatter(summary + source)

问题:大量 A+B 产物(观点提取/问题链)缺少 summary 和 source 字段。

修复方式:

# summary: 从正文第一个 # 标题提取
title_match = re.search(r'^#\s+(.+)', body, re.MULTILINE)
summary = title_match.group(1).strip()[:100] if title_match else md.stem

# source: 从文件路径推断
if '观点提取' in md.name or '问题链' in md.name:
    source = '素材库/播客/给女孩的商业第一课'

⚠️ PITFALL:GraphRAG 自动生成的实体文件(🧬 知识图谱/GraphRAG*/实体/)约 260+ 个没有 summary,这是正常的,不需要修复。

文件名含空格修复

# 扫描
find wiki/ -name "* *" -type f

# 重命名
mv "MarkItDown 操作手册.md" "MarkItDown-操作手册.md"

目录迁移提示页格式

当目录整体迁移时,旧位置保留 README.md 提示页:

---
title: "⚠️ 已迁移 - 目录名"
type: note
tags: [已废弃, 已迁移]
created: YYYY-MM-DD
summary: 已迁移到 ../../新位置。
source: "原系统名"
---

7. 事业线文件命名规范

事业线(50.团队/事业线/)的文件命名与 Agent 的对应关系:

文件名前缀Agent应放入子目录
O1-猎手 / 猎手-内容情报 / 猎手-渠道情报 / 猎手-技术调研 / 内容情报-给女孩的商业第一课猎手猎手/
O2-笔神笔神笔神/
O2-天眼天眼天眼/
O3-管家管家管家/
O4-极客极客极客/
O4-画神画神画神/
千金-千金千金/

特别注意:内容情报-给女孩的商业第一课-*.md 这些文件名(87个)不是系列名,是猎手的产出,需归入 猎手/ 目录。

O3-CFO 文件:属于 家庭线/CFO/,不是事业线。


8. 千金日报文件命名规范

文件名包含子目录
总裁日报总裁日报/
O1线/O2线/O3线/O4线O线日报/
晨间简报晨间简报/
浪漫-浪漫/
系统-系统/
其他其他/

反模式

  • ❌ 不要删除素材库/ 下的任何文件
  • ❌ Wiki 页面必须有 source 字段指向原始素材
  • ❌ 不要创建没有入链的孤立 wiki 页面(除索引页)
  • ❌ 清理前先 dry-run
  • ❌ 归档时不要改文件名
  • ✅ 每次大清理必须创建 归档清单-YYYY-MM-DD.md

相关 Skill

  • kb-ingest — 素材入库(详细流程)
  • kb-audit — 健康检查与清理(详细流程)
  • kb-tools — GBrain/GraphRAG/LLM Wiki 集成(详细配置)
  • chubby-knowledge-base-organization — 团队文件组织规范
  • content-ab-processing — A+B 双轨处理
  • memory-obsidian-sync — Memory 同步到 Obsidian
  • understand / understand-knowledge — 知识图谱可视化

相似的 Skill

pdf
anthropics/skills180k

pdf

Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating new PDFs, filling PDF forms, encrypting/decrypting PDFs, extracting images, and OCR on scanned PDFs to make them searchable. If the user mentions a .pdf file or asks to produce one, use this skill.

文档与办公

discernment-nudge
anthropics/skills180k

discernment-nudge

After you give a substantive answer or draft that the user may act on — advice or recommendations, drafted artifacts such as goals, plans, pitches, proposals, or emails, estimates or projections, analysis or interpretation of data, factual claims they may rely on, or a multi-step argument — invoke this skill BEFORE finalizing your reply and then, if it applies, append 2-3 short follow-up questions, each tied to something specific in what you just produced, that help the user check key facts, probe the reasoning or assumptions, and notice missing context. Do this at most once per conversation. Skip it when the user asked a trivial how-to or simple lookup, wants a purely educational explanation, asked you only to format, convert, or assemble a file from content they provided, is writing code they will run, is doing creative writing or casual chat, or already asked you to double-check, cite, or review — the skill file explains these boundaries and the exact output format.

文档与办公

doc-coauthoring
anthropics/skills180k

doc-coauthoring

Guide users through a structured workflow for co-authoring documentation. Use when user wants to write documentation, proposals, technical specs, decision docs, or similar structured content. This workflow helps users efficiently transfer context, refine content through iteration, and verify the doc works for readers. Trigger when user mentions writing docs, creating proposals, drafting specs, or similar documentation tasks.

文档与办公

docx
anthropics/skills180k

docx

Use this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx files) or Word templates (.dotx files). Triggers include: any mention of 'Word doc', 'word document', '.docx', '.dotx', or requests to produce professional documents with formatting like tables of contents, headings, page numbers, or letterheads. Also use when extracting or reorganizing content from .docx or .dotx files, inserting or replacing images in documents, performing find-and-replace in Word files, working with tracked changes or comments, or converting content into a polished Word document. If the user asks for a 'report', 'memo', 'letter', 'template', or similar deliverable as a Word or .docx file, use this skill. Do NOT use for PDFs, spreadsheets, Google Docs, or general coding tasks unrelated to document generation.

文档与办公

pptx
anthropics/skills180k

pptx

Use this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an email or summary); editing, modifying, or updating existing presentations; combining or splitting slide files; working with templates (.potx), layouts, speaker notes, or comments. Trigger whenever the user mentions "deck," "slides," "presentation," or references a .pptx or .potx filename, regardless of what they plan to do with the content afterward. If a .pptx or .potx file needs to be opened, created, or touched, use this skill.

文档与办公

canvas-design
anthropics/skills180k

canvas-design

Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.

文档与办公