Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

Flux Text to Speech (batch)

POST/v2/speakFlux Text to Speech (batch)

Synthesize a complete block of text into a single audio response using Deepgram's Flux TTS batch (REST) API. Use this for pre-rendering fixed audio (IVR prompts, notifications, narration) where the whole text is known up front and you don't need incremental playback or interruption.

Parameters

callbackstringquery

URL to which we'll make the callback request

callback_methodstringquery

HTTP method by which the callback request will be made

one of "POST", "PUT"

one of "POST", "PUT" · default "POST"

mip_opt_outbooleanquery

Opts out requests from the Deepgram Model Improvement Program. Refer to our Docs for pricing impacts before setting this to true. https://dpgr.am/deepgram-mip

default false

tagvaluequery

Label your requests for the purpose of identification during usage reporting

oneOf · 2 options
Option 1string
Option 2array of string
bit_ratevaluequery

The bitrate of the audio in bits per second. Choose from predefined ranges or specific values based on the encoding type.

oneOf · 3 options
Option 1stringV2SpeakPostParametersBitRate0

V2SpeakPostParametersBitRate0

Encoding - mp3(default). Supported bitrates - 8000, 16000, 24000, 32000, 40000, 48000(default) bps.

one of "8000", "16000", "24000", "32000", "40000", "48000"

Option 2integer

maximum 650000 · minimum 4000

Option 3integer

maximum 192000 · minimum 4000

containervaluequery

Container specifies the file format wrapper for the output audio. The available options depend on the encoding type.

oneOf · 5 options
Option 1stringV2SpeakPostParametersContainer0

V2SpeakPostParametersContainer0

No container.

one of "none"

Option 2stringV2SpeakPostParametersContainer1

V2SpeakPostParametersContainer1

Encoding - linear16. Supported container - wav (default), or no container.

one of "wav"

Option 3stringV2SpeakPostParametersContainer2

V2SpeakPostParametersContainer2

Encoding - mulaw. Supported container - wav (default), or no container.

one of "wav"

Option 4stringV2SpeakPostParametersContainer3

V2SpeakPostParametersContainer3

Encoding - alaw. Supported container - wav (default), or no container.

one of "wav"

Option 5stringV2SpeakPostParametersContainer4

V2SpeakPostParametersContainer4

Encoding - opus. Supported container - ogg (default).

one of "ogg"

encodingvaluequery

Encoding allows you to specify the expected encoding of your audio output

oneOf · 7 options
Option 1stringV2SpeakPostParametersEncoding0

V2SpeakPostParametersEncoding0

Encoding - linear16. Uncompressed, high-quality audio format often used for telephony or audio processing.

one of "linear16"

Option 2stringV2SpeakPostParametersEncoding1

V2SpeakPostParametersEncoding1

Encoding - flac. Lossless audio format for high-quality compression.

one of "flac"

Option 3stringV2SpeakPostParametersEncoding2

V2SpeakPostParametersEncoding2

Encoding - mulaw. Compressed audio format commonly used in telephony.

one of "mulaw"

Option 4stringV2SpeakPostParametersEncoding3

V2SpeakPostParametersEncoding3

Encoding - alaw. Similar to mulaw but used in international telephony.

one of "alaw"

Option 5stringV2SpeakPostParametersEncoding4

V2SpeakPostParametersEncoding4

Encoding - mp3. Popular compressed audio format for music and streaming.

one of "mp3"

Option 6stringV2SpeakPostParametersEncoding5

V2SpeakPostParametersEncoding5

Encoding - opus. High-compression audio format optimized for real-time communications.

one of "opus"

Option 7stringV2SpeakPostParametersEncoding6

V2SpeakPostParametersEncoding6

Encoding - aac. Advanced audio format offering better quality at smaller file sizes than mp3.

one of "aac"

expressivitystringquery

Expressive range of the generated speech, on a calm-to-animated axis. Accepted values: `-2`, `-1`, `0`, `1`, `2`. `0` (the default) is the voice's tuned delivery and the production-validated setting, with `-2` the calm end of the range and `2` the animated end. Supported on all Flux voices; applies to the whole request. Beta: behavior may change in future model versions, and non-default values increase the risk of hallucinations and pronunciation errors; audition before shipping. An invalid value is rejected with a `400` — `EXPRESSIVITY_OUT_OF_RANGE` for a value outside the range, `EXPRESSIVITY_INCREMENT_INVALID` for a fractional value. See [Expressivity](/docs/tts-expressivity).

one of "-2", "-1", "0", "1", "2"

one of "-2", "-1", "0", "1", "2"

modelstringqueryrequired

Flux TTS model used to synthesize the submitted text, in the form `flux-{voice}-{language}` (for example, `flux-alexis-en`). Required; unlike the v1 (Aura) endpoint there is no default and only flux models are accepted. English-only at launch.

sample_ratevaluequery

Sample Rate specifies the sample rate for the output audio. Based on the encoding, different sample rates are supported. For some encodings, the sample rate is not configurable

oneOf · 4 options
Option 1stringV2SpeakPostParametersSampleRate0

V2SpeakPostParametersSampleRate0

Encoding - linear16. Supported sample rates - 8000, 16000, 24000, 32000, 44100, 48000 Hz.

one of "8000", "16000", "24000", "32000", "44100", "48000"

Option 2stringV2SpeakPostParametersSampleRate1

V2SpeakPostParametersSampleRate1

Encoding - mulaw. Supported sample rates - 8000, 16000 Hz.

one of "8000", "16000"

Option 3stringV2SpeakPostParametersSampleRate2

V2SpeakPostParametersSampleRate2

Encoding - alaw. Supported sample rates - 8000, 16000 Hz.

one of "8000", "16000"

Option 4stringV2SpeakPostParametersSampleRate3

V2SpeakPostParametersSampleRate3

Encoding - flac. Supported sample rates - 8000, 16000, 22050, 32000, 48000 Hz.

one of "8000", "16000", "22050", "32000", "48000"

speednumber · doublequery

Speaking rate multiplier that adjusts the pace of generated speech while preserving natural prosody and voice quality. Accepted values run `0.5` to `1.5` in `0.05` increments. Not yet supported in all languages.

default 1 · maximum 1.5 · minimum 0.5 · multipleOf 0.05

prioritystringquery

Processing priority for asynchronous (callback) requests. The only supported value is low.

one of "low"

one of "low"

Request body

Transform text to speech

application/json
objectSpeakV2Request

SpeakV2Request

Request body for Flux TTS batch (REST) text-to-speech conversion. The full block of text is synthesized in a single request and returned as one audio response.

textstringrequired

The text content to be converted to speech. The server normalizes and preprocesses the text before synthesis. Inline pause and pronunciation controls are not yet applied; they are stripped from the text before synthesis.

Example request
{
  "text": "string"
}

Responses

200Returns the synthesized audio in the requested encoding as a binary stream. When a `callback` URL is supplied, the request is processed asynchronously and the response body is instead a JSON acknowledgement (Content-Type `application/json`) of the form {"request_id": "..."}, with the audio delivered to the callback URL. Because this endpoint is typed as a binary audio stream, SDK callers that set `callback` receive this JSON acknowledgement through the audio byte iterator as raw bytes and must join the chunks and parse `request_id` themselves.application/json
objectSpeakV2AcceptedResponse

SpeakV2AcceptedResponse

Accepted response returned when a callback URL is supplied; the audio is delivered asynchronously to that URL.

request_idstring · uuidrequired

Unique identifier for tracking the asynchronous request

Example response
{
  "request_id": "00000000-0000-0000-0000-000000000000"
}
400Invalid Request. Inline pause and pronunciation controls are not applied and are stripped rather than rejected.application/json
valueErrorResponse

ErrorResponse

oneOf · 3 options
Option 1stringErrorResponseTextError

ErrorResponseTextError

Option 2objectErrorResponseLegacyError

ErrorResponseLegacyError

err_codestring

The error code

err_msgstring

The error message

request_idstring

The request ID

Option 3objectErrorResponseModernError

ErrorResponseModernError

categorystring

The category of the error

detailsstring

A description of the error

messagestring

A message about the error

request_idstring

The unique identifier of the request

Example response
{
  "err_code": "string",
  "err_msg": "string",
  "request_id": "string"
}
Documentation menu