Skip to content
FunCoding

Search

Search docs, Skills and MCP

podcast-transcribe

播客/小宇宙 → 下载 → 转录 → 存为 Markdown 的完整工作流。 支持 RSS 批量下载、单集链接转录。

文档与办公1.2kpodcast-transcribe/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/chubbyguan/chubbyskills/podcast-transcribe/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

播客转录 Skill

将播客音频下载并转录为 Markdown。支持小宇宙、喜马拉雅、直接音频地址、本地音频和 RSS 批量流程。

默认在本地使用 SenseVoice-Small(与视频类技能共用的 chubby_common/funasr.py 封装)。可选云端后端为阿里云百炼 DashScope 的 qwen3-asr-flash 和 Groq 的 whisper-large-v3-turbo,用户明确选择后才启用;音频会发送至云端并可能计费。

环境要求

以下命令在本 skill 目录运行,建议使用 Python 3.11 或更新版本:

python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -r requirements.txt
# macOS: brew install ffmpeg
# Ubuntu: sudo apt install ffmpeg

本地模型首次使用时需要下载。仅使用云端后端不需要 funasr 本地依赖;不支持的音频容器转换仍可能需要 ffmpeg。云端密钥通过安全环境配置,不能写入命令参数、转录稿或版本库。

单集与批量

python3 scripts/transcribe.py "https://www.xiaoyuzhoufm.com/episode/xxxxx" ./output \
  --provider local
python3 scripts/transcribe.py "/你的音频目录/episode.mp3" ./output \
  --provider local --language zh
python3 scripts/batch_transcribe.py --rss-url "替换为实际 RSS 地址" \
  --output ./output --count 10 --provider local

将示例地址替换为实际来源。单集页面会尝试提取音频链接,提取失败时改用直接音频地址或本地文件。标准输出最后一行是成功生成的 Markdown 路径。

自动下载仅接受公网 HTTP(S) 直连地址,禁用代理和重定向,拒绝本地/私网地址。需要跳转或代理的来源,请先自行下载音频,再传入本地文件路径。

本地模型选择:SenseVoice-Small vs Qwen3-ASR-0.6B

--provider local 支持两个本地模型:SenseVoiceSmall(默认)和 qwen3-asr-0.6b:

python3 scripts/transcribe.py "/你的音频目录/episode.mp3" ./output --provider local
python3 scripts/transcribe.py "/你的音频目录/episode.mp3" ./output --provider local --model qwen3-asr-0.6b

qwen3-asr-0.6b 是可选重依赖,不在默认安装内,需自行 pip install qwen-asr transformers torch。

以下对比数据来自 2026-10 在 MacBook Pro(Apple M3 Pro,纯 CPU)上对 5 分钟中文播客的实测:

SenseVoice-Small(默认)Qwen3-ASR-0.6B(可选)
速度RTF 0.11(5 分钟音频约 33 秒)RTF 0.80(约 4 分钟,慢约 7 倍)
专有名词一般("岩茶"误作"盐茶",人名前后不一致)更稳("岩茶"、人名识别一致)
主要风险输出混入情感标签,正式文稿需清洗有幻觉式改写风险("黄金加工厂"→"皇帝家族");无 ITN,数字输出为全文字
长音频VAD 自动分段,稳定整段进模型;本后端已把 max_new_tokens 调到 4096 避免截断
适用场景CPU 默认选择,长播客友好GPU 机器,或对专名/人名准确性敏感的内容

两者质量互有胜负、没有代差。CPU 场景请保持默认 SenseVoice-Small;有 GPU 或专名敏感时再选 Qwen3-ASR-0.6B。

可选云端转录

后端凭据环境变量默认模型
local无SenseVoiceSmall(仅用于元数据记录)
dashscopeDASHSCOPE_API_KEYqwen3-asr-flash
groqGROQ_API_KEYwhisper-large-v3-turbo

配置 DASHSCOPE_API_KEY 或 GROQ_API_KEY 后:

python3 scripts/transcribe.py "/你的音频目录/episode.mp3" ./output \
  --provider dashscope --cloud-timeout 1800 \
  --state-dir "$HOME/.local/state/chubbyskills/podcast"
python3 scripts/transcribe.py "/你的音频目录/episode.mp3" ./output \
  --provider groq --cloud-timeout 1800 \
  --state-dir "$HOME/.local/state/chubbyskills/podcast"

--provider 优先于 PODCAST_TRANSCRIBE_PROVIDER,都未指定时使用 local。批量入口支持同样的 provider、模型、语言、等待和状态目录参数,并把最终选项显式传给单集进程。

两个云端后端都是同步接口。DashScope 限制为编码后不超过 10MB、时长不超过 5 分钟(客户端在原始文件超过 7 MiB 时拒绝并提示改用本地 SenseVoice)。Groq 免费层单文件上限 25MB,达到上限的长音频会自动分片:ffmpeg 切成 20 分钟一段(约 9.6MB/段),逐段转录后按顺序拼接,时间戳自动累加偏移;分片进度逐段落盘,中断后重跑同一命令断点续传,不重复提交已完成分片;遇 429 按 Retry-After/指数退避等待。Groq 返回的分段时间戳会作为附录保留。分片和容器转换需要本机 ffmpeg。

云端任务恢复

默认状态目录为 ~/.local/state/chubbyskills/podcast;设置了 XDG_STATE_HOME 时使用其中的 chubbyskills/podcast。还可用 CHUBBY_PODCAST_STATE_DIR 或 --state-dir 指定。目录包含任务与转录内容,应按音频内容的隐私要求保存。

相同音频、provider、模型、语言和服务地址再次运行时,已有任务继续查询;已完成结果可以重新导出。提交结果不明确或服务端报告失败时,不自动重新提交。普通网络重试保留状态目录并重跑原命令即可。

--resubmit 明确创建新任务,可能重复计费。不要用它解决单纯的轮询超时。删除状态目录或改变输入配置也可能失去复用条件。客户端超时不等于服务端取消,也不代表没有计费。

完整仓库使用说明、统一入库和验证范围见云端转录说明。

产物与限制

产物为带 frontmatter、来源和转录后端标记的 Markdown。后端返回可用分段时保留时间戳;没有时间信息时不编造时间轴。

  • 本地 SenseVoice-Small 只输出纯文本全文,不生成逐段时间戳(段数记为 1)。
  • 本地 CPU 推理耗时受音频长度和机器配置影响。
  • 转录可能有专有名词、数字或断句错误,引用前核对原始音频。
  • 云端真实可用性、账号权限、音频兼容性和费用尚需独立验收。
  • 本 skill 不提供通用说话人分离保证。

贡献与参考

可选云端转录需求最初来自 binyangzhu000-sudo 的 PR #3 和 Anil-matcha 的 PR #5(Atlas / MuAPI 实验后端,现已被 DashScope 后端取代,归属保留)。

合规声明

请遵守来源平台条款并尊重内容版权,控制请求频率。云端处理前确认自己有权向所选服务提交音频;下载和转录不会改变原内容的版权归属。

Similar Skills

pdf
anthropics/skills180k

pdf

Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating new PDFs, filling PDF forms, encrypting/decrypting PDFs, extracting images, and OCR on scanned PDFs to make them searchable. If the user mentions a .pdf file or asks to produce one, use this skill.

Docs & office

discernment-nudge
anthropics/skills180k

discernment-nudge

After you give a substantive answer or draft that the user may act on — advice or recommendations, drafted artifacts such as goals, plans, pitches, proposals, or emails, estimates or projections, analysis or interpretation of data, factual claims they may rely on, or a multi-step argument — invoke this skill BEFORE finalizing your reply and then, if it applies, append 2-3 short follow-up questions, each tied to something specific in what you just produced, that help the user check key facts, probe the reasoning or assumptions, and notice missing context. Do this at most once per conversation. Skip it when the user asked a trivial how-to or simple lookup, wants a purely educational explanation, asked you only to format, convert, or assemble a file from content they provided, is writing code they will run, is doing creative writing or casual chat, or already asked you to double-check, cite, or review — the skill file explains these boundaries and the exact output format.

Docs & office

doc-coauthoring
anthropics/skills180k

doc-coauthoring

Guide users through a structured workflow for co-authoring documentation. Use when user wants to write documentation, proposals, technical specs, decision docs, or similar structured content. This workflow helps users efficiently transfer context, refine content through iteration, and verify the doc works for readers. Trigger when user mentions writing docs, creating proposals, drafting specs, or similar documentation tasks.

Docs & office

docx
anthropics/skills180k

docx

Use this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx files) or Word templates (.dotx files). Triggers include: any mention of 'Word doc', 'word document', '.docx', '.dotx', or requests to produce professional documents with formatting like tables of contents, headings, page numbers, or letterheads. Also use when extracting or reorganizing content from .docx or .dotx files, inserting or replacing images in documents, performing find-and-replace in Word files, working with tracked changes or comments, or converting content into a polished Word document. If the user asks for a 'report', 'memo', 'letter', 'template', or similar deliverable as a Word or .docx file, use this skill. Do NOT use for PDFs, spreadsheets, Google Docs, or general coding tasks unrelated to document generation.

Docs & office

pptx
anthropics/skills180k

pptx

Use this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an email or summary); editing, modifying, or updating existing presentations; combining or splitting slide files; working with templates (.potx), layouts, speaker notes, or comments. Trigger whenever the user mentions "deck," "slides," "presentation," or references a .pptx or .potx filename, regardless of what they plan to do with the content afterward. If a .pptx or .potx file needs to be opened, created, or touched, use this skill.

Docs & office

canvas-design
anthropics/skills180k

canvas-design

Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.

Docs & office