/v1/speakText to Speech transformationConvert text into natural-sounding speech using Deepgram's TTS REST API
Parameters
callbackstringqueryURL to which we'll make the callback request
callback_methodstringqueryHTTP method by which the callback request will be made
mip_opt_outbooleanqueryOpts out requests from the Deepgram Model Improvement Program. Refer to our Docs for pricing impacts before setting this to true. https://dpgr.am/deepgram-mip
tagvaluequeryLabel your requests for the purpose of identification during usage reporting
oneOf · 2 options
bit_ratevaluequeryThe bitrate of the audio in bits per second. Choose from predefined ranges or specific values based on the encoding type.
oneOf · 3 options
V1SpeakPostParametersBitRate0
Encoding - mp3(default). Supported bitrates - 32000, 48000(default) bps.
containervaluequeryContainer specifies the file format wrapper for the output audio. The available options depend on the encoding type.
oneOf · 5 options
V1SpeakPostParametersContainer0
No container.
V1SpeakPostParametersContainer1
Encoding - linear16. Supported container - wav (default), or no container.
V1SpeakPostParametersContainer2
Encoding - mulaw. Supported container - wav (default), or no container.
V1SpeakPostParametersContainer3
Encoding - alaw. Supported container - wav (default), or no container.
V1SpeakPostParametersContainer4
Encoding - opus. Supported container - ogg (default).
encodingvaluequeryEncoding allows you to specify the expected encoding of your audio output
oneOf · 7 options
V1SpeakPostParametersEncoding0
Encoding - linear16. Uncompressed, high-quality audio format often used for telephony or audio processing.
V1SpeakPostParametersEncoding1
Encoding - flac. Lossless audio format for high-quality compression.
V1SpeakPostParametersEncoding2
Encoding - mulaw. Compressed audio format commonly used in telephony.
V1SpeakPostParametersEncoding3
Encoding - alaw. Similar to mulaw but used in international telephony.
V1SpeakPostParametersEncoding4
Encoding - mp3. Popular compressed audio format for music and streaming.
V1SpeakPostParametersEncoding5
Encoding - opus. High-compression audio format optimized for real-time communications.
V1SpeakPostParametersEncoding6
Encoding - aac. Advanced audio format offering better quality at smaller file sizes than mp3.
modelstringqueryAI model used to process submitted text
sample_ratevaluequerySample Rate specifies the sample rate for the output audio. Based on the encoding, different sample rates are supported. For some encodings, the sample rate is not configurable
oneOf · 5 options
V1SpeakPostParametersSampleRate0
Encoding - linear16. Supported sample rates - 8000, 16000, 24000, 32000, 48000 Hz.
V1SpeakPostParametersSampleRate1
Encoding - mulaw. Supported sample rates - 8000, 16000 Hz.
V1SpeakPostParametersSampleRate2
Encoding - alaw. Supported sample rates - 8000, 16000 Hz.
V1SpeakPostParametersSampleRate3
Encoding - mp3. Sample rate is fixed and not configurable (22050 Hz).
V1SpeakPostParametersSampleRate4
Encoding - opus. Sample rate is fixed at 48000 Hz.
speednumber · doublequerySpeaking rate multiplier that adjusts the pace of generated speech while preserving natural prosody and voice quality. Not yet supported in all languages.
Request body
Transform text to speech
application/json
SpeakV1Request
Request body for text-to-speech conversion
textstringrequiredThe text content to be converted to speech
{
"text": "string"
}Responses
speak_v1_audio_generate_Response_200
Empty response body
{}ErrorResponse
oneOf · 3 options
ErrorResponseTextError
ErrorResponseLegacyError
err_codestringThe error code
err_msgstringThe error message
request_idstringThe request ID
ErrorResponseModernError
categorystringThe category of the error
detailsstringA description of the error
messagestringA message about the error
request_idstringThe unique identifier of the request
{
"err_code": "string",
"err_msg": "string",
"request_id": "string"
}