/v2/speakFlux Text to Speech (batch)Synthesize a complete block of text into a single audio response using Deepgram's Flux TTS batch (REST) API. Use this for pre-rendering fixed audio (IVR prompts, notifications, narration) where the whole text is known up front and you don't need incremental playback or interruption.
Parameters
callbackstringqueryURL to which we'll make the callback request
callback_methodstringqueryHTTP method by which the callback request will be made
mip_opt_outbooleanqueryOpts out requests from the Deepgram Model Improvement Program. Refer to our Docs for pricing impacts before setting this to true. https://dpgr.am/deepgram-mip
tagvaluequeryLabel your requests for the purpose of identification during usage reporting
oneOf · 2 options
bit_ratevaluequeryThe bitrate of the audio in bits per second. Choose from predefined ranges or specific values based on the encoding type.
oneOf · 3 options
V2SpeakPostParametersBitRate0
Encoding - mp3(default). Supported bitrates - 8000, 16000, 24000, 32000, 40000, 48000(default) bps.
containervaluequeryContainer specifies the file format wrapper for the output audio. The available options depend on the encoding type.
oneOf · 5 options
V2SpeakPostParametersContainer0
No container.
V2SpeakPostParametersContainer1
Encoding - linear16. Supported container - wav (default), or no container.
V2SpeakPostParametersContainer2
Encoding - mulaw. Supported container - wav (default), or no container.
V2SpeakPostParametersContainer3
Encoding - alaw. Supported container - wav (default), or no container.
V2SpeakPostParametersContainer4
Encoding - opus. Supported container - ogg (default).
encodingvaluequeryEncoding allows you to specify the expected encoding of your audio output
oneOf · 7 options
V2SpeakPostParametersEncoding0
Encoding - linear16. Uncompressed, high-quality audio format often used for telephony or audio processing.
V2SpeakPostParametersEncoding1
Encoding - flac. Lossless audio format for high-quality compression.
V2SpeakPostParametersEncoding2
Encoding - mulaw. Compressed audio format commonly used in telephony.
V2SpeakPostParametersEncoding3
Encoding - alaw. Similar to mulaw but used in international telephony.
V2SpeakPostParametersEncoding4
Encoding - mp3. Popular compressed audio format for music and streaming.
V2SpeakPostParametersEncoding5
Encoding - opus. High-compression audio format optimized for real-time communications.
V2SpeakPostParametersEncoding6
Encoding - aac. Advanced audio format offering better quality at smaller file sizes than mp3.
expressivitystringqueryExpressive range of the generated speech, on a calm-to-animated axis. Accepted values: `-2`, `-1`, `0`, `1`, `2`. `0` (the default) is the voice's tuned delivery and the production-validated setting, with `-2` the calm end of the range and `2` the animated end. Supported on all Flux voices; applies to the whole request. Beta: behavior may change in future model versions, and non-default values increase the risk of hallucinations and pronunciation errors; audition before shipping. An invalid value is rejected with a `400` — `EXPRESSIVITY_OUT_OF_RANGE` for a value outside the range, `EXPRESSIVITY_INCREMENT_INVALID` for a fractional value. See [Expressivity](/docs/tts-expressivity).
modelstringqueryrequiredFlux TTS model used to synthesize the submitted text, in the form `flux-{voice}-{language}` (for example, `flux-alexis-en`). Required; unlike the v1 (Aura) endpoint there is no default and only flux models are accepted. English-only at launch.
sample_ratevaluequerySample Rate specifies the sample rate for the output audio. Based on the encoding, different sample rates are supported. For some encodings, the sample rate is not configurable
oneOf · 4 options
V2SpeakPostParametersSampleRate0
Encoding - linear16. Supported sample rates - 8000, 16000, 24000, 32000, 44100, 48000 Hz.
V2SpeakPostParametersSampleRate1
Encoding - mulaw. Supported sample rates - 8000, 16000 Hz.
V2SpeakPostParametersSampleRate2
Encoding - alaw. Supported sample rates - 8000, 16000 Hz.
V2SpeakPostParametersSampleRate3
Encoding - flac. Supported sample rates - 8000, 16000, 22050, 32000, 48000 Hz.
speednumber · doublequerySpeaking rate multiplier that adjusts the pace of generated speech while preserving natural prosody and voice quality. Accepted values run `0.5` to `1.5` in `0.05` increments. Not yet supported in all languages.
prioritystringqueryProcessing priority for asynchronous (callback) requests. The only supported value is low.
Request body
Transform text to speech
application/json
SpeakV2Request
Request body for Flux TTS batch (REST) text-to-speech conversion. The full block of text is synthesized in a single request and returned as one audio response.
textstringrequiredThe text content to be converted to speech. The server normalizes and preprocesses the text before synthesis. Inline pause and pronunciation controls are not yet applied; they are stripped from the text before synthesis.
{
"text": "string"
}Responses
SpeakV2AcceptedResponse
Accepted response returned when a callback URL is supplied; the audio is delivered asynchronously to that URL.
request_idstring · uuidrequiredUnique identifier for tracking the asynchronous request
{
"request_id": "00000000-0000-0000-0000-000000000000"
}ErrorResponse
oneOf · 3 options
ErrorResponseTextError
ErrorResponseLegacyError
err_codestringThe error code
err_msgstringThe error message
request_idstringThe request ID
ErrorResponseModernError
categorystringThe category of the error
detailsstringA description of the error
messagestringA message about the error
request_idstringThe unique identifier of the request
{
"err_code": "string",
"err_msg": "string",
"request_id": "string"
}