Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Cross-Turn Context

On /v1/speak, every request is independent: text goes in, audio comes out, and all state is discarded. Short responses like “Of course” lose the tone established earlier in the conversation, and prosody can shift at chunk boundaries. /v2/speak closes this gap at the model layer by persisting conversational state across turns — with no new API parameters to set or manage.

This page explains the mechanism. The protocol provides the surface (Client Messages, Server Messages); this is what happens underneath.

Flux TTS maintains internal state that evolves as it generates. Cross-turn context works by not resetting that state between turns:

  1. Turn 1 — State starts from the conditioning clip only. The model generates; state now includes the clip plus the generated audio.
  2. The user responds — handled by your STT and LLM, not by TTS.
  3. Turn 2 — The model generates from the existing state. Delivery is informed by the prior output, so prosody and pacing carry forward.
  4. Turn N — State has accumulated across the conversation. There is no prompt re-consumption and no replay, which is a meaningful cost and latency advantage over re-priming each turn.

Ending a turn does not reset the model’s conversational state — it carries forward into the next turn, which is what keeps prosody consistent across a back-and-forth.

Action Resets model state
Flush (ends the turn) No
New connection Yes — each session starts fresh
  • State persistence adds no per-turn compute. The only cost is GPU memory held for the session duration.
  • Max session duration is 1 hour; the server closes the WebSocket at that mark. See Feature Overview → Session Limits.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu