Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Voice Agent Integration Patterns

Flux TTS is built to sit downstream of an LLM in a voice agent pipeline. This guide shows the common integration patterns, pairing Flux TTS’s /v2/speak with a streaming STT (such as Flux STT) and an LLM. Each pattern is intentionally small — drop it into your pipeline and adapt.

For the messages used here, see Client Messages and Server Messages.

On each end-of-turn from STT, stream the LLM response into the active turn and flush when the response is complete.

Python
async def agent_loop(stt_conn, speak_conn, llm):
    async for event in stt_conn:
        if event.event == "EndOfTurn":
            async for token in llm.generate_stream(event.transcript):
                await speak_conn.send({"type": "Speak", "text": token})
            await speak_conn.send({"type": "Flush"})

When the user speaks over the agent, stop local playback immediately, then send Interrupt. The server returns text_spoken / text_remaining so you can reconcile LLM context with exactly what the user heard. The split is computed from an optional playback position you include — how much of the turn’s audio actually played before the barge-in — so it reflects what the user really heard, not just what the server sent.

Python
async def handle_barge_in(stt_event, speak_conn, playback):
    if stt_event.event == "StartOfTurn":
        playback.stop()  # stop audio locally, immediately
        await speak_conn.send({
            "type": "Interrupt",
            # ms of audio played since the START OF THE SESSION (not the turn)
            "playback_offset": {"type": "time_ms", "value": playback.session_offset_ms()}
        })
        # On SpeechInterrupted: llm_context.append(assistant=text_spoken)

See Interruption Handling for the full pattern, including in-flight audio and edge cases.

Streaming text into Flux TTS is mostly about not doing extra work — your LLM plumbing almost certainly handles it already:

  1. Keep the whitespace your LLM emits. Most LLM token streams already include the spaces between tokens; just don’t strip them. The one thing to confirm: when you stitch together two separate generations (a reply, then a tool-call result, then another reply), keep a space between them — otherwise "Hello world." followed by "How are you?" becomes "Hello world.How are you?".
  2. Don’t chunk or buffer text. You don’t need to detect sentence boundaries or hold text back — the server streams a turn’s audio as tokens arrive. Send them as they come and Flush only when the agent’s response is complete.

The Flux TTS WebSocket is the framework-to-Deepgram leg of your pipeline, not the end-user audio leg. For the user-facing real-time audio path, voice agent frameworks typically use WebRTC. Keep that distinction in mind when reasoning about end-to-end latency: the TTS WebSocket carries the framework-to-Deepgram leg, while stopping audio in the user’s ear is always your client’s job, done locally.


Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu