Voice Agent Integration Patterns
Flux TTS is built to sit downstream of an LLM in a voice agent pipeline. This guide shows the common integration patterns, pairing Flux TTS’s /v2/speak with a streaming STT (such as Flux STT) and an LLM. Each pattern is intentionally small — drop it into your pipeline and adapt.
For the messages used here, see Client Messages and Server Messages.
Pattern 1: Basic agent loop
Section titled “Pattern 1: Basic agent loop”On each end-of-turn from STT, stream the LLM response into the active turn and flush when the response is complete.
async def agent_loop(stt_conn, speak_conn, llm):
async for event in stt_conn:
if event.event == "EndOfTurn":
async for token in llm.generate_stream(event.transcript):
await speak_conn.send({"type": "Speak", "text": token})
await speak_conn.send({"type": "Flush"})Pattern 2: Barge-in
Section titled “Pattern 2: Barge-in”When the user speaks over the agent, stop local playback immediately, then send Interrupt. The server returns text_spoken / text_remaining so you can reconcile LLM context with exactly what the user heard. The split is computed from an optional playback position you include — how much of the turn’s audio actually played before the barge-in — so it reflects what the user really heard, not just what the server sent.
async def handle_barge_in(stt_event, speak_conn, playback):
if stt_event.event == "StartOfTurn":
playback.stop() # stop audio locally, immediately
await speak_conn.send({
"type": "Interrupt",
# ms of audio played since the START OF THE SESSION (not the turn)
"playback_offset": {"type": "time_ms", "value": playback.session_offset_ms()}
})
# On SpeechInterrupted: llm_context.append(assistant=text_spoken)See Interruption Handling for the full pattern, including in-flight audio and edge cases.
Streaming text correctly
Section titled “Streaming text correctly”Streaming text into Flux TTS is mostly about not doing extra work — your LLM plumbing almost certainly handles it already:
- Keep the whitespace your LLM emits. Most LLM token streams already include the spaces between tokens; just don’t strip them. The one thing to confirm: when you stitch together two separate generations (a reply, then a tool-call result, then another reply), keep a space between them — otherwise
"Hello world."followed by"How are you?"becomes"Hello world.How are you?". - Don’t chunk or buffer text. You don’t need to detect sentence boundaries or hold text back — the server streams a turn’s audio as tokens arrive. Send them as they come and
Flushonly when the agent’s response is complete.
A note on transport
Section titled “A note on transport”The Flux TTS WebSocket is the framework-to-Deepgram leg of your pipeline, not the end-user audio leg. For the user-facing real-time audio path, voice agent frameworks typically use WebRTC. Keep that distinction in mind when reasoning about end-to-end latency: the TTS WebSocket carries the framework-to-Deepgram leg, while stopping audio in the user’s ear is always your client’s job, done locally.
Related resources
Section titled “Related resources”- The Speech Lifecycle — the state machine behind these loops
- Client Messages / Server Messages — full wire reference
- Cross-Turn Context — what persists across turns
- Build a Flux-enabled Voice Agent (STT) — the STT side of the pipeline