The Speech Lifecycle and State Machine
Flux TTS reframes synthesis from a text-to-audio pipe into a turn-based conversation. You own turn boundaries (Flush) and content; the server handles streaming and lifecycle reporting. This page explains the state machine that connects your Client Messages to the Server Messages you receive.
Vocabulary
Section titled “Vocabulary”Getting the unit right is the key to reasoning about the protocol:
- Chunk — the text payload of one
Speakmessage. Size is up to you: a single LLM token or a full paragraph. Chunk boundaries do not drive synthesis. - Turn — one complete agent response, bounded by
Flush. The turn is the reporting unit:SpeechStartedfires once at the start, andSpeechMetadatafires once at the end. speech_id— the server-assigned identifier for a turn (dg_sp_<12 hex>). Informational; a new one is minted at the start of each turn.
The server holds one active turn at a time. If you Flush and then send more Speak before the active turn finishes, the new turn is pending — pending turns queue behind the active one (there’s no limit) and become active in order.
You stream chunks; the server groups them into a turn; the wire reports per turn.
States
Section titled “States”| State | Meaning |
|---|---|
Idle |
Connection open, no active turn. The initial state after Connected. |
Generating |
A turn is active: the server accepts Speak for it and streams its audio as text arrives. |
Finalizing |
You’ve sent Flush; the server is finishing the active turn’s remaining audio. A Speak sent now starts a pending turn. |
Closing |
Close received; finishing queued audio, emitting SessionMetadata, then closing the socket. |
Key rules
Section titled “Key rules”- The first
Speak(fromIdle) starts a turn. The server assigns aspeech_idand emitsSpeechStarted.speech_idis turn-scoped — it represents one agent turn. - Subsequent
Speakmessages append to the active turn. The server streams that turn’s audio as text arrives — you don’t have toFlushto start hearing audio. SpeechStartedandSpeechMetadatabookmark a turn’s audio. Every audio frame for a turn arrives between them, andSpeechMetadatais the server’s signal that no more audio is coming for that turn.- Manual
Flushfinalizes the active turn: the server finishes its remaining audio, emitsFlushedwhen the buffer has actually been flushed, thenSpeechMetadata. - One active turn at a time. A
Speaksent after youFlush(but before that turn’sSpeechMetadata) starts a pending turn; it becomes active once the current turn’sSpeechMetadatais sent, and gets its ownSpeechStarted. SessionMetadatais sent once before close, with cumulative totals.
Edge cases
Section titled “Edge cases”| Event during state | Behavior |
|---|---|
Speak during Finalizing (after Flush, before SpeechMetadata) |
Starts a pending turn; it becomes active after the current turn’s SpeechMetadata. |
Flush with no active turn |
No-op; the server emits a NO_ACTIVE_SPEECH warning and the connection stays open. |
A turn, end to end
Section titled “A turn, end to end”A single agent turn that completes normally — the agent says “Sure, I can help you cancel your subscription,” streamed as LLM tokens and ended with a manual Flush. The exchange is shown as alternating client and server steps.
1. The client streams tokens. The first Speak starts the turn.
{"type": "Speak", "text": "Sure, "}
{"type": "Speak", "text": "I can help you "}
{"type": "Speak", "text": "cancel your subscription."}2. The server opens the turn and streams audio.
{"type": "SpeechStarted", "speech_id": "dg_sp_a1b2c3d4e5f6"}
// binary audio frames stream as generation proceeds3. The client ends the turn.
{"type": "Flush"}4. The server finishes the audio and reports the turn. All of the turn’s audio has arrived between SpeechStarted and this SpeechMetadata.
{"type": "Flushed", "speech_id": "dg_sp_a1b2c3d4e5f6"}
// remaining binary audio frames
{
"type": "SpeechMetadata",
"speech_id": "dg_sp_a1b2c3d4e5f6",
"audio_duration_ms": 3200,
"input_character_count": 47,
"billable_character_count": 47,
"controls_applied": {
"pronunciations_applied": 0,
"breaks_applied": 0,
"pronunciation_warnings": 0
}
}The next Speak begins a new turn with a new speech_id.
Related resources
Section titled “Related resources”- Client Messages — the messages that drive these transitions
- Server Messages — full reference for each event above
- Build a Flux TTS Voice Agent — the state machine in a real agent loop
- Cross-Turn Context — what persists across turns