Streaming TTS Internals
Sentence chunker, streaming provider ABC, capability matrix and how to add a streaming TTS provider
Hermes can stream TTS audio as it arrives from the provider, instead of waiting
for the full audio before playing. This is used by voice mode (CLI/TUI live
conversation), the dashboard speak-stream WebSocket, and — via the gateway
StreamingTTSConsumer — any platform adapter that opts into streaming audio.
Voice replies start speaking after the first clause instead of after full
generation + synthesis.
Architecture
The streaming pipeline has four parts:
- Producer — the LLM emits text deltas as it generates a response
- Sentence chunker —
tools.tts_streaming.SentenceChunkeraccumulates deltas, strips<think>blocks (even split across deltas), and flushes complete sentences - TTS provider — a registered
StreamingTTSProviderturns each sentence into raw PCM chunks (int16 mono at the provider's declaredsample_rate) - Audio sink —
sounddevice.OutputStreamfor local playback (tools.tts_tool_speaker.stream_tts_to_speaker), or a gateway platform adapter'swrite_streaming_ttsseam (gateway/streaming_tts_consumer.py)
Providers with no chunked API still get per-sentence playback via the proven
sync text_to_speech_tool path, so edge (the default) is conversational too.
All spoken text is cleaned by tools.tts_text_normalize.prepare_spoken_text
(one cleaner, all paths).
How to pick a provider
By default the dispatcher streams with the provider you already configured
(tts.provider) when that provider has a chunked API — it never silently
swaps your voice for a different provider just to get streaming.
To override, set tts.streaming.provider in your config.yaml:
- a provider name (
elevenlabs,gemini,openai,xai) pins that streamer autowalks the priority listelevenlabs → gemini → openai → xaiand uses the first one whose credentials resolve — an explicit opt-in to "best chunked voice available"
tts:
provider: gemini
streaming:
provider: gemini # or "auto"
min_len: 20 # shortest first sentence (chars) spoken on its own; CJK setups use ~6
gemini:
model: gemini-2.5-flash-preview-tts
voice: KoreCapability matrix
| Provider | Transport | Chunked PCM | Credentials |
|---|---|---|---|
| elevenlabs | chunked HTTP (pcm_24000) | yes | ELEVENLABS_API_KEY / tts.elevenlabs |
| openai | chunked HTTP (with_streaming_response, pcm) | yes | tts.openai.api_key → env → managed gateway |
| gemini | SSE (streamGenerateContent?alt=sse) | yes | GEMINI_API_KEY / GOOGLE_API_KEY |
| xai | WebSocket (wss://api.x.ai/v1/tts) | yes | XAI_API_KEY preferred, else xAI OAuth (the subscription bearer 403s on metered TTS) |
| edge, piper, kitten, neutts, mistral, minimax, deepinfra, … | — | no (per-sentence sync fallback) | as usual |
plugin TTSProvider with streams_pcm | the plugin's stream(format="pcm") | yes, at its stream_sample_rate | the plugin's own |
All credential lookups go through resolve_provider_secret()
(config > env/.env > credential pool) — never bare env reads. Streamed bodies
are capped at 16 MiB per sentence, mirroring the sync providers' bounded
upstream-body invariant.
Adding a new streaming provider
- Subclass
StreamingTTSProviderintools/tts_streaming.py - Set
sample_rate(andchannels/sample_widthif not int16 mono) - Implement
available()(a pure probe — never install anything) andstream(self, text) -> Iterator[bytes]yielding raw PCM chunks - Decorate with
@register("yourname") - Add tests in
tests/tools/test_tts_streaming.py
The ABC enforces the contract; the registry makes the provider discoverable;
the dispatcher (stream_tts_to_speaker) and the gateway consumer handle the
sentence buffer, stop events, and audio sink for free.
Plugin providers
A plugin TTSProvider (registered with ctx.register_tts_provider()) joins this
path without core changes. It sets streams_pcm = True and a positive
stream_sample_rate, and its stream(text, format="pcm", voice=, model=, speed=)
yields int16 mono PCM. resolve_streaming_provider adapts it
(tools.tts_streaming._PluginPCMStreamer) under the same routing rules as the
sync dispatcher (tools.tts_tool_plugins): built-in names never reach the
registry, a same-named type: command provider wins, and the plugin gets the
same tts.voice / tts.model / tts.speed as synthesize(). A configured
plugin streams first. Under tts.streaming.provider: auto it is tried only after
the built-in priority list. streams_pcm, stream_sample_rate and
is_available() are read every time a streamer is resolved. If the rate is
missing or the provider reports unavailable, Hermes keeps per-sentence synthesis.
As with the built-in streamers, the speaker pipeline prefetches up to three
sentences, so a plugin's stream() must tolerate concurrent calls (the sync
path serializes synthesize(); this path does not).
Gateway streaming (platform adapters)
gateway/streaming_tts_consumer.py bridges agent deltas to an adapter's
streaming-audio seam. Adapters opt in by overriding, on
BasePlatformAdapter:
supports_streaming_tts(chat_id, audio_format) -> boolbegin_streaming_tts / write_streaming_tts / finish_streaming_tts / abort_streaming_tts
All default to unsupported/no-op, so existing adapters are untouched. When a turn's streaming audio completes, the whole-file auto-TTS reply for that turn is suppressed (no double playback); when streaming fails before any audio was audible, the gateway falls back to the legacy whole-file voice reply.