跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

scenario-audio

Use when generating or handling audio on Scenario via MCP. Triggers include music tracks, full-length songs with vocals written from lyrics, background scores, soundtracks, game sound effects, SFX, foley, ambience, looping audio, voiceover, narration, speech, TTS, text-to-speech, dialogue, voice cloning, re-voicing a recording, scoring or adding sound to a video, transcription, or requests to create, wait on, play, or download audio files (MP3, WAV) with Scenario tools.

AI 与智能体923skills/scenario-audio/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/scenario-labs/skills/scenario-audio/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Scenario Audio Generation

Overview

Scenario generates audio through the same loop as images. The live catalog covers three generation lanes (music, sound effects, voice/speech) plus video-to-audio soundtrack models and audio utilities. Connection and the core generation loop: see the scenario skill. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Quick reference

StepToolNotes
Find a modelrecommend with the need in the user's own words; search only for a member known by namecapability="txt2audio" covers music, SFX, and TTS; optional, inferred from the prompt when omitted
Inspect inputsmodel_schema_getaudio schemas vary widely: durations, lyrics, voices, looping
Generatemodel_runschema-conformant parameters; wait=false for long jobs
Waitjobs_waitblocks server-side; on timeout re-call with pending_job_ids
Listenasset_displayrenders an inline audio player
Saveasset_downloadreturns a download URL: curl -L -o out.mp3 "<url>"

Find existing audio assets with search target="assets", filters={kind: "audio"}. Team and project scope (team_id, project_id): see the scenario skill.

What the audio surface covers

  • Music: text-to-music models produce short beds or full-length songs with vocals; the song lane has its own contract, below.
  • Sound effects: text-to-SFX models generate short clips from a description; some support seamless looping.
  • Voice and speech: text-to-speech with preset voices, multilingual output, and emotion or pacing controls; some clone a voice from a short clip, and speech-to-speech re-voices a recording.
  • Video to audio: models that score a silent video or add synchronized effects.
  • Utilities: model_scenario-audio-cut, model_scenario-audio-split, model_scenario-audio-extract, and model_scenario-compose-video (fixed ids: each is Scenario's single deterministic tool for its operation, so discovery would only re-derive them); the compositor lays a finished track (score, voiceover, re-voiced take) over a clip as an audio layer, per scenario-video-assembly; for speech-to-text transcription, recommend with the need in the user's own words.
  • Stem separation: one named stem per run (discover with recommend), vocals included, with no instrumental option. Voice isolation returns the clean speech and never the removed music and effects as a second stem, so a two-stem split (voice against everything else) is a gap to report, not a member to keep hunting for.

Per-family contracts: scenario-elevenlabs (speech, dubbing, re-voicing, music, SFX), scenario-ace-step and scenario-minimax-music (songs), scenario-sonilo (SFX and video scoring).

Worked example: a game sound effect

  1. recommend with capability="txt2audio" and the user's own words as prompt ("a game sound effect: a heavy wooden chest creaking open"). The ranking returns txt2audio models such as model_elevenlabs-sound-effects-v2 (example only).
  2. model_schema_get with that model_id. Returns the exact fields: prompt plus controls such as duration or looping.
  3. model_run with the same model_id and parameters={"prompt": "heavy wooden treasure chest creaking open, single event, dry, no music"}.
  4. jobs_wait job_ids=["job_xxx"] on any job_id returned without assets (in_progress after a timed-out wait, the backend's queued or in-progress after wait=false), re-calling with the returned pending_job_ids on timeout.
  5. asset_display asset_id="asset_xxx" to play it inline.
  6. asset_download with no format, then save the returned URL with curl -L.

Prompting tips:

  • SFX: name the source, material, action, and acoustic space, and say what to exclude ("no music", "no reverb"). One event per clip; generate variations as separate runs.
  • Music: give genre, mood, tempo, and instrumentation. Short beds usually take a single prompt, with duration or looping in the schema.
  • Speech: keep the text field to the words to speak; voice, language, emotion, and pacing live in separate schema fields or inline tags.

Speech and dialogue

Discover a voice member with recommend (capability="txt2audio", the user's words, "two-person dialogue" when it is one), then read its text field's description in model_schema_get: that is where a member usually states its own delivery grammar.

  • Tag syntax is per member, even within one family: square brackets ([whispers]), angle brackets (<sigh>, <short pause>), parentheses ((sighs)), or wrapping pairs (<whisper>text</whisper>), and some read none. A tag in the wrong grammar can be spoken aloud, so copy the spelling from the text field's description; where it names none, from the member's catalog description or recommend's notes, and with no source write no tags. A correctly spelled tag is a request, not a guarantee: two identical runs have disagreed on one, so have the user listen before a take is final and re-run a take whose tag was skipped.
  • Two voices, one take. A member with a multi-speaker array (up to 2 rows of a speaker label and a voice at authoring time) reads the text as turns, one Name: line per turn, each Name matching a row's label exactly; an unprefixed line continues the previous turn, and the single-voice field is ignored in that mode. Preset voice names say nothing about gender, age, or tone, so state which preset plays whom and let the user confirm before the paid run. A scene with more speakers is split into takes of at most two, delivered in order or laid over the picture per scenario-video-assembly.
  • Text is capped and priced. The text field carries a max_length (5000 characters on one member at authoring time), an overrun is a 400, and cost_impact marks it as the price driver: split a long script at turn boundaries and dry_run the first take.
  • Pin the language through the schema's language field when there is one, rather than naming it in the text.

Songs with vocals

A full-length song is not a longer music bed, and song schemas vary more than the rest of the lane, so model_schema_get decides the shape: a style prompt plus a separate lyric sheet, one prose prompt carrying both, or an ordered section array with per-section text and styles.

  • Words never go in a style field. Where the schema splits the two, the style field carries genre, mood, tempo, key, vocal style, and instrumentation; the lyric field carries the words, shaped by section tags such as [Verse] and [Chorus].
  • Instrumental and auto-lyrics are flags where the schema has them; asking for either in prose is unreliable, and where no flag exists the text fields are the only lever. Flipping the instrumental flag on a rerun gives a different take, not the same song: seed, where a model has one, only repeats identical settings. For an instrumental of a track you already have, try an audio2audio cover model (discover with recommend), checking the schema since not all carry the flag.
  • Text fields are length-capped per model and field, and going over is a 400 rather than a truncation.

Where the schema exposes a duration field (flagged cost_impact), it caps both length and price; where none exists, the lyric sheet or prompt sets both. Either way, price the song with dry_run: true before committing, then launch with wait: false; both are model_run arguments, not parameters keys.

Repeatability and batching are per member, not per lane: at authoring time the repaint members took numOutputs (1 to 4) and no seed, so a repaint that must keep the same singer is run as a batch and picked from, while the section composers had seed and no numOutputs. Read both off model_schema_get before promising either.

Extending a song

Making an existing track longer is its own audio2audio lane, not a longer text-to-music run: recommend with capability="audio2audio" and the extension need in the user's words; use search only when the member is already known by name. The member that does it takes the song as audio and an ordered sections array (up to 30 at authoring time) where each entry either keeps a slice of the original (sourceStartSeconds and sourceEndSeconds) or generates a new one (text with [Verse]-style tags, durationSeconds, positiveStyles as an array), with contextAdherence deciding how closely new sections follow their neighbors. Kept slices bill like generated audio of the same length, so dry_run the whole plan first. The repaint, edit and add-layer members regenerate inside the original's duration and never lengthen it, and recommend ranked an older text-to-music member for "extend a song" at authoring time: an extension is not a txt2audio need, so do not take that pick.

Common mistakes

  • Hardcoding generative model IDs: availability differs per team and evolves. Re-discover each session, recommend for the need or search for a name; only the fixed first-party tool ids above stay constant.
  • Skipping model_schema_get: one audio model's parameters will not fit another (voices, durations, and lyric fields all differ).
  • Polling job_get in a loop: music jobs can run minutes. Use jobs_wait; on timeout re-call with pending_job_ids.
  • Pasting raw asset URLs into chat: use asset_display to play audio.
  • Passing format to asset_download for audio: it converts image formats only, so omit it.
  • Putting voice direction inside TTS text ("say this angrily"): direction can end up spoken. Use the schema's emotion or voice fields.
  • Writing a dialogue as one voice reading both parts: on a multi-speaker member, fill the speaker rows and prefix every turn with its label, since with the rows empty the member reads everything in its single voice.
  • Pasting lyrics into the style field: the model then describes a song instead of singing one.
  • Answering "make it longer" with a repaint or a new text-to-music run: the first keeps the duration, the second loses the song; the extend lane above keeps the slices the user chose.
  • Putting dry_run or wait inside parameters: they are model_run's own arguments, so a stray dry_run still charges and a stray wait blocks up to 180s.

相似的 Skill

brand-guidelines
anthropics/skills180k

brand-guidelines

Applies Anthropic's official brand colors and typography to any sort of artifact that may benefit from having Anthropic's look-and-feel. Use it when brand colors or style guidelines, visual formatting, or company design standards apply.

AI 与智能体

internal-comms
anthropics/skills180k

internal-comms

A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use this skill whenever asked to write some sort of internal communications (status reports, leadership updates, 3P updates, company newsletters, FAQs, incident reports, project updates, etc.).

AI 与智能体

template-skill
anthropics/skills180k

template-skill

Replace with description of the skill and when Claude should use it.

AI 与智能体

mcp-builder
anthropics/skills180k

mcp-builder

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

AI 与智能体

algorithmic-art
anthropics/skills180k

algorithmic-art

Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.

AI 与智能体

academy-guide
anthropics/skills180k

academy-guide

Stop and check this skill before finishing any reply to a question about how to use Claude or a Claude product — it recommends matching courses, tutorials, and use cases from Claude Academy (academy.claude.com), Anthropic's learning hub. Trigger on: "how do I", "how can I", "getting started with", "what can Claude do", "teach me", "learn to use"; questions about artifacts, projects, skills, plugins, connectors, MCP; requests about rolling Claude out to a team, class, or organization; and any ask for training materials, onboarding content, or learning resources. Use it when the user is learning how to use a feature or product — not when they are mid-task and just want the task done. This skill composes with other skills: after consulting product documentation to answer how a Claude feature works, also check here for a matching course or tutorial — a docs-grounded answer and an Academy recommendation belong together. Only recommend on a strong match; never invent Academy content.

AI 与智能体