跳到正文
FunCoding

搜索

搜索文档、文章、Skill 和 MCP

Text-to-speech quickstart

Turn on text-to-speech, pick a provider, and send a first audio reply

Quick start

Pick a provider

OpenAI and ElevenLabs are the most reliable hosted options. Microsoft and Local CLI work without an API key. See the provider matrix for the full list.

Set the API key

Export the env var for your provider (for example OPENAI_API_KEY, ELEVENLABS_API_KEY). Microsoft and Local CLI need no key.

Enable in config

Set tts.auto: "always" and tts.provider:

{
  tts: {
    auto: "always",
    provider: "elevenlabs",
  },
}

Try it in chat

/tts status shows the current state. /tts audio Hello from OpenClaw sends a one-off audio reply.

Auto-TTS is off by default. When tts.provider is unset, OpenClaw picks the first configured provider in registry auto-select order. The built-in tts agent tool is explicit-intent only: ordinary chat stays text unless the user asks for audio, uses /tts, or enables Auto-TTS/directive speech.

Feishu and WhatsApp voice notes need ffmpeg on the Gateway host when the channel must convert the provider's audio to Ogg/Opus. Already-compatible audio skips this conversion. If conversion fails, Feishu sends the original audio as a file attachment; the WhatsApp send fails. See TTS output for the transcoding rules.

Supported providers

ProviderAuthNotes
Azure SpeechAZURE_SPEECH_KEY + AZURE_SPEECH_REGION (also AZURE_SPEECH_API_KEY, SPEECH_KEY, SPEECH_REGION)Native Ogg/Opus voice-note output and telephony.
DeepInfraDEEPINFRA_API_KEYOpenAI-compatible TTS. Defaults to hexgrad/Kokoro-82M.
ElevenLabsELEVENLABS_API_KEY or XI_API_KEYVoice cloning, multilingual, deterministic via seed; streamed for Discord voice playback.
Fish AudioFISH_API_KEY or FISH_AUDIO_API_KEYS2.1 hosted TTS, expressive tags, voice discovery, streaming, and telephony.
Google GeminiGEMINI_API_KEY or GOOGLE_API_KEYGemini API TTS; opt in to Gemini 3.8, where persona style is metadata, not spoken text.
GradiumGRADIUM_API_KEYVoice-note and telephony output.
InworldINWORLD_API_KEYStreaming TTS API. Native Opus voice-note and PCM telephony.
Local CLInoneRuns a configured local TTS command.
MicrosoftnonePublic Edge neural TTS via node-edge-tts. Best-effort, no SLA.
MiniMaxMINIMAX_API_KEY (or Token Plan: MINIMAX_OAUTH_TOKEN, MINIMAX_CODE_PLAN_KEY, MINIMAX_CODING_API_KEY)T2A v2 API. Defaults to speech-2.8-hd.
OpenAIOPENAI_API_KEYAlso used for auto-summary; supports persona instructions.
OpenRouterOPENROUTER_API_KEY (can reuse models.providers.openrouter.apiKey)Default model hexgrad/kokoro-82m.
VolcengineVOLCENGINE_TTS_API_KEY or BYTEPLUS_SEED_SPEECH_API_KEY (legacy AppID/token: VOLCENGINE_TTS_APPID/_TOKEN)BytePlus Seed Speech HTTP API.
VydraVYDRA_API_KEYShared image, video, and speech provider.
xAIXAI_API_KEYxAI batch TTS. Native Opus voice-note is not supported.
Xiaomi MiMoXIAOMI_API_KEYMiMo TTS through Xiaomi chat completions.

If multiple providers are configured, the selected one is used first and the others are fallback options. Auto-summary uses summaryModel (or agents.defaults.model.primary), so that provider must also be authenticated if you keep summaries enabled.

The bundled Microsoft provider uses Microsoft Edge's online neural TTS service via node-edge-tts. It is a public web service without a published SLA or quota — treat it as best-effort. The legacy provider id edge is normalized to microsoft and openclaw doctor --fix rewrites persisted config; new configs should always use microsoft.