/v1/listenTranscribe and analyze pre-recorded audio and videoTranscribe audio and video using Deepgram's speech-to-text REST API
Parameters
callbackstringqueryURL to which we'll make the callback request
callback_methodstringqueryHTTP method by which the callback request will be made
extravaluequeryArbitrary key-value pairs that are attached to the API response for usage in downstream processing
oneOf · 2 options
sentimentbooleanqueryRecognizes the sentiment throughout a transcript or text
summarizevaluequerySummarize content. For Listen API, supports string version option. For Read API, accepts boolean only.
oneOf · 2 options
V1ListenPostParametersSummarize0
tagvaluequeryLabel your requests for the purpose of identification during usage reporting
oneOf · 2 options
topicsbooleanqueryDetect topics throughout a transcript or text
custom_topicvaluequeryCustom topics you want the model to detect within your input audio or text if present Submit up to `100`.
oneOf · 2 options
custom_topic_modestringquerySets how the model will interpret strings submitted to the `custom_topic` param. When `strict`, the model will only return topics submitted using the `custom_topic` param. When `extended`, the model will return its own detected topics in addition to those submitted using the `custom_topic` param
intentsbooleanqueryRecognizes speaker intent throughout a transcript or text
custom_intentvaluequeryCustom intents you want the model to detect within your input audio if present
oneOf · 2 options
custom_intent_modestringquerySets how the model will interpret intents submitted to the `custom_intent` param. When `strict`, the model will only return intents submitted using the `custom_intent` param. When `extended`, the model will return its own detected intents in the `custom_intent` param.
detect_entitiesbooleanqueryIdentifies and extracts key entities from content in submitted audio
detect_languagevaluequeryIdentifies the dominant language spoken in submitted audio
oneOf · 2 options
diarizebooleanqueryDeprecated: use `diarize_model` instead. Recognize speaker changes. Each word in the transcript will be assigned a speaker number starting at 0.
diarize_modelstringquerySelect and enable a specific diarization model version. Specifying this parameter enables diarization and selects the model — you do not need to also set the deprecated `diarize=true` parameter. For batch, supported values are `latest` (currently v2), `v1`, and `v2`. For streaming, supported values are `latest` (currently v1) and `v1`; `v2` returns a validation error on streaming requests.
dictationbooleanqueryDictation mode for controlling formatting with dictated speech
encodingstringquerySpecify the expected encoding of your submitted audio
filler_wordsbooleanqueryFiller Words can help transcribe interruptions in your audio, like "uh" and "um"
keytermarrayqueryKey term prompting improves recognition of specialized terminology and brands. Only compatible with Nova-3. `keyterm` accepts plain terms only. Unlike the legacy `keywords` feature, it does not support weights or intensifiers. Appending one (for example, `keyterm=term:0.15`) is not rejected—the weight is silently ignored and the entire value is treated as a literal keyterm. To boost multiple separate keyterms, repeat the `keyterm` parameter (for example, `keyterm=term1&keyterm=term2`). To boost one multi-word phrase as a single keyterm, join the words with `%20` or `+` (for example, `keyterm=customer%20service`). Do not separate keyterms with commas, semicolons, or line breaks.
keywordsvaluequeryKeywords can boost or suppress specialized terminology and brands. `keywords` is not supported with Nova-3 models; use `keyterm` instead.
oneOf · 2 options
languagestringqueryThe [BCP-47 language tag](https://tools.ietf.org/html/bcp47) that hints at the primary spoken language. Depending on the Model and API endpoint you choose only certain languages are available
measurementsbooleanquerySpoken measurements will be converted to their corresponding abbreviations
modelvaluequeryAI model used to process submitted audio
oneOf · 2 options
V1ListenPostParametersModel0
Our public models available to all accounts
multichannelbooleanqueryTranscribe each audio channel independently
numeralsbooleanqueryNumerals converts numbers from written format to numerical format
paragraphsbooleanquerySplits audio into paragraphs to improve transcript readability
profanity_filterbooleanqueryProfanity Filter looks for recognized profanity and converts it to the nearest recognized non-profane word or removes it from the transcript completely
punctuatebooleanqueryAdd punctuation and capitalization to the transcript
redactvaluequeryRedaction removes sensitive information from your transcripts
oneOf · 2 options
V1ListenPostParametersRedact1
replacevaluequerySearch for terms or phrases in submitted audio and replaces them
oneOf · 2 options
searchvaluequerySearch for terms or phrases in submitted audio
oneOf · 2 options
smart_formatbooleanqueryApply formatting to transcript output. When set to true, additional formatting will be applied to transcripts to improve readability
utterancesbooleanquerySegments speech into meaningful semantic units
utt_splitnumber · doublequerySeconds to wait before detecting a pause between words in submitted audio
versionvaluequeryVersion of an AI model to use
oneOf · 2 options
V1ListenPostParametersVersion0
Use the latest version of a model
mip_opt_outbooleanqueryOpts out requests from the Deepgram Model Improvement Program. Refer to our Docs for pricing impacts before setting this to true. https://dpgr.am/deepgram-mip
Request body
Transcribe an audio or video file
application/json
ListenV1RequestUrl
Audio file URL to transcribe
urlstring · urirequired{
"url": "https://example.com"
}Responses
listen_v1_media_transcribe_Response_200
oneOf · 2 options
ListenV1Response
The standard transcription response
metadataobjectrequiredShow child attributes
channelsintegerrequiredcreatedstring · date-timerequireddiarize_infoobjectThe diarizer that produced the speaker labels. Present only when a diarizer ran.
Show child attributes
archstringrequiredThe diarizer arch, such as `v1` or `v2`
model_uuidstringrequiredThe diarizer model UUID
durationnumber · doublerequiredintents_infoobjectShow child attributes
input_tokensintegermodel_uuidstringoutput_tokensintegermodel_infoobjectrequiredmodelsarray of stringrequiredShow child attributes
request_idstring · uuidrequiredsentiment_infoobjectShow child attributes
input_tokensintegermodel_uuidstringoutput_tokensintegersha256stringrequiredsummary_infoobjectShow child attributes
input_tokensintegermodel_uuidstringoutput_tokensintegertagsarray of stringShow child attributes
topics_infoobjectShow child attributes
input_tokensintegermodel_uuidstringoutput_tokensintegertransaction_keystringresultsobjectrequiredShow child attributes
channelsarray of objectrequiredShow child attributes
Show array items
alternativesarray of objectShow child attributes
Show array items
confidencestringentitiesarray of objectShow child attributes
Show array items
confidencestringend_wordstringlabelstringraw_valuestringstart_wordstringvaluestringparagraphsobjectShow child attributes
paragraphsarray of objectShow child attributes
Show array items
endstringnum_wordsintegersentencesarray of objectShow child attributes
Show array items
endstringstartstringtextstringspeakerintegerstartstringtranscriptstringsummariesarray of objectShow child attributes
Show array items
end_wordstringstart_wordstringsummarystringtopicsarray of objectShow child attributes
Show array items
end_wordstringstart_wordstringtextstringtopicsarray of stringShow child attributes
transcriptstringwordsarray of objectShow child attributes
Show array items
confidencestringendstringspeakerintegerThe speaker of the word, present when diarization is enabled
speaker_confidencestringConfidence in the speaker assignment. Returned only for pre-recorded diarization; not available for streaming
startstringwordstringdetected_languagestringsearcharray of objectShow child attributes
Show array items
hitsarray of objectShow child attributes
Show array items
confidencestringendstringsnippetstringstartstringquerystringintentsobjectOutput whenever `intents=true` is used
sentimentsobjectOutput whenever `sentiment=true` is used
summaryobjectShow child attributes
resultstringshortstringtopicsobjectOutput whenever `topics=true` is used
utterancesarray of objectShow child attributes
Show array items
channelintegerconfidencestringendstringidstring · uuidspeakerintegerstartstringtranscriptstringwordsarray of objectShow child attributes
Show array items
confidencestringendstringpunctuated_wordstringspeakerintegerspeaker_confidencestringstartstringwordstringListenV1AcceptedResponse
Accepted response for asynchronous transcription requests
request_idstring · uuidrequiredUnique identifier for tracking the asynchronous request
{
"request_id": "00000000-0000-0000-0000-000000000000"
}ListenV1Response
The standard transcription response
{
"metadata": {
"channels": 0,
"created": "2026-06-09T00:00:00Z",
"diarize_info": {
"arch": "string",
"model_uuid": "string"
},
"duration": 0,
"intents_info": {
"input_tokens": 0,
"model_uuid": "string",
"output_tokens": 0
},
"model_info": {},
"models": [
"string"
],
"request_id": "00000000-0000-0000-0000-000000000000",
"sentiment_info": {
"input_tokens": 0,
"model_uuid": "string",
"output_tokens": 0
},
"sha256": "string",
"summary_info": {
"input_tokens": 0,
"model_uuid": "string",
"output_tokens": 0
},
"tags": [
"string"
],
"topics_info": {
"input_tokens": 0,
"model_uuid": "string",
"output_tokens": 0
},
"transaction_key": "deprecated"
},
"results": {
"channels": [
{
"alternatives": [
{
"confidence": "string",
"entities": [
{
"confidence": "string",
"end_word": "string",
"label": "string",
"raw_value": "string",
"start_word": "string",
"value": "string"
}
],
"paragraphs": {
"paragraphs": [
{
"end": "string",
"num_words": 0,
"sentences": [],
"speaker": 0,
"start": "string"
}
],
"transcript": "string"
},
"summaries": [
{
"end_word": "string",
"start_word": "string",
"summary": "string"
}
],
"topics": [
{
"end_word": "string",
"start_word": "string",
"text": "string",
"topics": [
"string"
]
}
],
"transcript": "string",
"words": [
{
"confidence": "string",
"end": "string",
"speaker": 0,
"speaker_confidence": "string",
"start": "string",
"word": "string"
}
]
}
],
"detected_language": "string",
"search": [
{
"hits": [
{
"confidence": "string",
"end": "string",
"snippet": "string",
"start": "string"
}
],
"query": "string"
}
]
}
],
"intents": {
"segments": [
{
"end_word": 0
}
]
}
}
}