Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

Transcribe and analyze pre-recorded audio and video

POST/v1/listenTranscribe and analyze pre-recorded audio and video

Transcribe audio and video using Deepgram's speech-to-text REST API

Parameters

callbackstringquery

URL to which we'll make the callback request

callback_methodstringquery

HTTP method by which the callback request will be made

one of "POST", "PUT"

one of "POST", "PUT" · default "POST"

extravaluequery

Arbitrary key-value pairs that are attached to the API response for usage in downstream processing

oneOf · 2 options
Option 1string
Option 2array of string
sentimentbooleanquery

Recognizes the sentiment throughout a transcript or text

default false

summarizevaluequery

Summarize content. For Listen API, supports string version option. For Read API, accepts boolean only.

oneOf · 2 options
Option 1stringV1ListenPostParametersSummarize0

V1ListenPostParametersSummarize0

one of "v2"

Option 2boolean

default false

tagvaluequery

Label your requests for the purpose of identification during usage reporting

oneOf · 2 options
Option 1string
Option 2array of string
topicsbooleanquery

Detect topics throughout a transcript or text

default false

custom_topicvaluequery

Custom topics you want the model to detect within your input audio or text if present Submit up to `100`.

oneOf · 2 options
Option 1string
Option 2array of string
custom_topic_modestringquery

Sets how the model will interpret strings submitted to the `custom_topic` param. When `strict`, the model will only return topics submitted using the `custom_topic` param. When `extended`, the model will return its own detected topics in addition to those submitted using the `custom_topic` param

one of "extended", "strict"

one of "extended", "strict" · default "extended"

intentsbooleanquery

Recognizes speaker intent throughout a transcript or text

default false

custom_intentvaluequery

Custom intents you want the model to detect within your input audio if present

oneOf · 2 options
Option 1string
Option 2array of string
custom_intent_modestringquery

Sets how the model will interpret intents submitted to the `custom_intent` param. When `strict`, the model will only return intents submitted using the `custom_intent` param. When `extended`, the model will return its own detected intents in the `custom_intent` param.

one of "extended", "strict"

one of "extended", "strict" · default "extended"

detect_entitiesbooleanquery

Identifies and extracts key entities from content in submitted audio

default false

detect_languagevaluequery

Identifies the dominant language spoken in submitted audio

oneOf · 2 options
Option 1boolean

default false

Option 2array of string
diarizebooleanquery

Deprecated: use `diarize_model` instead. Recognize speaker changes. Each word in the transcript will be assigned a speaker number starting at 0.

default false

diarize_modelstringquery

Select and enable a specific diarization model version. Specifying this parameter enables diarization and selects the model — you do not need to also set the deprecated `diarize=true` parameter. For batch, supported values are `latest` (currently v2), `v1`, and `v2`. For streaming, supported values are `latest` (currently v1) and `v1`; `v2` returns a validation error on streaming requests.

one of "latest", "v1", "v2"

one of "latest", "v1", "v2"

dictationbooleanquery

Dictation mode for controlling formatting with dictated speech

default false

encodingstringquery

Specify the expected encoding of your submitted audio

one of "linear16", "flac", "mulaw", "amr-nb", "amr-wb", "opus", "speex", "g729"

one of "linear16", "flac", "mulaw", "amr-nb", "amr-wb", "opus", "speex", "g729"

filler_wordsbooleanquery

Filler Words can help transcribe interruptions in your audio, like "uh" and "um"

default false

keytermarrayquery

Key term prompting improves recognition of specialized terminology and brands. Only compatible with Nova-3. `keyterm` accepts plain terms only. Unlike the legacy `keywords` feature, it does not support weights or intensifiers. Appending one (for example, `keyterm=term:0.15`) is not rejected—the weight is silently ignored and the entire value is treated as a literal keyterm. To boost multiple separate keyterms, repeat the `keyterm` parameter (for example, `keyterm=term1&keyterm=term2`). To boost one multi-word phrase as a single keyterm, join the words with `%20` or `+` (for example, `keyterm=customer%20service`). Do not separate keyterms with commas, semicolons, or line breaks.

keywordsvaluequery

Keywords can boost or suppress specialized terminology and brands. `keywords` is not supported with Nova-3 models; use `keyterm` instead.

oneOf · 2 options
Option 1string
Option 2array of string
languagestringquery

The [BCP-47 language tag](https://tools.ietf.org/html/bcp47) that hints at the primary spoken language. Depending on the Model and API endpoint you choose only certain languages are available

default "en"

measurementsbooleanquery

Spoken measurements will be converted to their corresponding abbreviations

default false

modelvaluequery

AI model used to process submitted audio

oneOf · 2 options
Option 1stringV1ListenPostParametersModel0

V1ListenPostParametersModel0

Our public models available to all accounts

one of "nova-3", "nova-3-general", "nova-3-medical", "nova-2", "nova-2-general", "nova-2-meeting", "nova-2-finance", "nova-2-conversationalai", "nova-2-voicemail", "nova-2-video", "nova-2-medical", "nova-2-drivethru", "nova-2-automotive", "nova", "nova-general", "nova-phonecall", "nova-medical", "enhanced", "enhanced-general", "enhanced-meeting", "enhanced-phonecall", "enhanced-finance", "base", "meeting", "phonecall", "finance", "conversationalai", "voicemail", "video"

Option 2string
multichannelbooleanquery

Transcribe each audio channel independently

default false

numeralsbooleanquery

Numerals converts numbers from written format to numerical format

default false

paragraphsbooleanquery

Splits audio into paragraphs to improve transcript readability

default false

profanity_filterbooleanquery

Profanity Filter looks for recognized profanity and converts it to the nearest recognized non-profane word or removes it from the transcript completely

default false

punctuatebooleanquery

Add punctuation and capitalization to the transcript

default false

redactvaluequery

Redaction removes sensitive information from your transcripts

oneOf · 2 options
Option 1string
Option 2array of stringV1ListenPostParametersRedact1

V1ListenPostParametersRedact1

replacevaluequery

Search for terms or phrases in submitted audio and replaces them

oneOf · 2 options
Option 1string
Option 2array of string
searchvaluequery

Search for terms or phrases in submitted audio

oneOf · 2 options
Option 1string
Option 2array of string
smart_formatbooleanquery

Apply formatting to transcript output. When set to true, additional formatting will be applied to transcripts to improve readability

default false

utterancesbooleanquery

Segments speech into meaningful semantic units

default false

utt_splitnumber · doublequery

Seconds to wait before detecting a pause between words in submitted audio

default 0.8

versionvaluequery

Version of an AI model to use

oneOf · 2 options
Option 1stringV1ListenPostParametersVersion0

V1ListenPostParametersVersion0

Use the latest version of a model

one of "latest"

Option 2string
mip_opt_outbooleanquery

Opts out requests from the Deepgram Model Improvement Program. Refer to our Docs for pricing impacts before setting this to true. https://dpgr.am/deepgram-mip

default false

Request body

Transcribe an audio or video file

application/json
objectListenV1RequestUrl

ListenV1RequestUrl

Audio file URL to transcribe

urlstring · urirequired
Example request
{
  "url": "https://example.com"
}

Responses

200Returns either transcription results, or a request_id when using a callback.application/json
valuelisten_v1_media_transcribe_Response_200

listen_v1_media_transcribe_Response_200

oneOf · 2 options
Option 1objectListenV1Response

ListenV1Response

The standard transcription response

metadataobjectrequired
Show child attributes
channelsintegerrequired
createdstring · date-timerequired
diarize_infoobject

The diarizer that produced the speaker labels. Present only when a diarizer ran.

Show child attributes
archstringrequired

The diarizer arch, such as `v1` or `v2`

model_uuidstringrequired

The diarizer model UUID

durationnumber · doublerequired
intents_infoobject
Show child attributes
input_tokensinteger
model_uuidstring
output_tokensinteger
model_infoobjectrequired
modelsarray of stringrequired
Show child attributes
request_idstring · uuidrequired
sentiment_infoobject
Show child attributes
input_tokensinteger
model_uuidstring
output_tokensinteger
sha256stringrequired
summary_infoobject
Show child attributes
input_tokensinteger
model_uuidstring
output_tokensinteger
tagsarray of string
Show child attributes
topics_infoobject
Show child attributes
input_tokensinteger
model_uuidstring
output_tokensinteger
transaction_keystring

default "deprecated"

resultsobjectrequired
Show child attributes
channelsarray of objectrequired
Show child attributes
Show array items
alternativesarray of object
Show child attributes
Show array items
confidencestring
entitiesarray of object
Show child attributes
Show array items
confidencestring
end_wordstring
labelstring
raw_valuestring
start_wordstring
valuestring
paragraphsobject
Show child attributes
paragraphsarray of object
Show child attributes
Show array items
endstring
num_wordsinteger
sentencesarray of object
Show child attributes
Show array items
endstring
startstring
textstring
speakerinteger
startstring
transcriptstring
summariesarray of object
Show child attributes
Show array items
end_wordstring
start_wordstring
summarystring
topicsarray of object
Show child attributes
Show array items
end_wordstring
start_wordstring
textstring
topicsarray of string
Show child attributes
transcriptstring
wordsarray of object
Show child attributes
Show array items
confidencestring
endstring
speakerinteger

The speaker of the word, present when diarization is enabled

speaker_confidencestring

Confidence in the speaker assignment. Returned only for pre-recorded diarization; not available for streaming

startstring
wordstring
detected_languagestring
searcharray of object
Show child attributes
Show array items
hitsarray of object
Show child attributes
Show array items
confidencestring
endstring
snippetstring
startstring
querystring
intentsobject

Output whenever `intents=true` is used

Show child attributes
segmentsarray of object
Show child attributes
Show array items
end_wordnumber · double
intentsarray of object
Show child attributes
Show array items
confidence_scorestring
intentstring
start_wordnumber · double
textstring
sentimentsobject

Output whenever `sentiment=true` is used

Show child attributes
averageobject
Show child attributes
sentimentstring
sentiment_scorenumber · double
segmentsarray of object
Show child attributes
Show array items
end_wordnumber · double
sentimentstring
sentiment_scorenumber · double
start_wordnumber · double
textstring
summaryobject
Show child attributes
resultstring
shortstring
topicsobject

Output whenever `topics=true` is used

Show child attributes
segmentsarray of object
Show child attributes
Show array items
end_wordnumber · double
start_wordnumber · double
textstring
topicsarray of object
Show child attributes
Show array items
confidence_scorestring
topicstring
utterancesarray of object
Show child attributes
Show array items
channelinteger
confidencestring
endstring
idstring · uuid
speakerinteger
startstring
transcriptstring
wordsarray of object
Show child attributes
Show array items
confidencestring
endstring
punctuated_wordstring
speakerinteger
speaker_confidencestring
startstring
wordstring
Option 2objectListenV1AcceptedResponse

ListenV1AcceptedResponse

Accepted response for asynchronous transcription requests

request_idstring · uuidrequired

Unique identifier for tracking the asynchronous request

Example response
{
  "request_id": "00000000-0000-0000-0000-000000000000"
}
400Invalid Requestapplication/json
objectListenV1Response

ListenV1Response

The standard transcription response

metadataobjectrequiredListenV1ResponseMetadata ↑
resultsobjectrequiredListenV1ResponseResults ↑
Example response
{
  "metadata": {
    "channels": 0,
    "created": "2026-06-09T00:00:00Z",
    "diarize_info": {
      "arch": "string",
      "model_uuid": "string"
    },
    "duration": 0,
    "intents_info": {
      "input_tokens": 0,
      "model_uuid": "string",
      "output_tokens": 0
    },
    "model_info": {},
    "models": [
      "string"
    ],
    "request_id": "00000000-0000-0000-0000-000000000000",
    "sentiment_info": {
      "input_tokens": 0,
      "model_uuid": "string",
      "output_tokens": 0
    },
    "sha256": "string",
    "summary_info": {
      "input_tokens": 0,
      "model_uuid": "string",
      "output_tokens": 0
    },
    "tags": [
      "string"
    ],
    "topics_info": {
      "input_tokens": 0,
      "model_uuid": "string",
      "output_tokens": 0
    },
    "transaction_key": "deprecated"
  },
  "results": {
    "channels": [
      {
        "alternatives": [
          {
            "confidence": "string",
            "entities": [
              {
                "confidence": "string",
                "end_word": "string",
                "label": "string",
                "raw_value": "string",
                "start_word": "string",
                "value": "string"
              }
            ],
            "paragraphs": {
              "paragraphs": [
                {
                  "end": "string",
                  "num_words": 0,
                  "sentences": [],
                  "speaker": 0,
                  "start": "string"
                }
              ],
              "transcript": "string"
            },
            "summaries": [
              {
                "end_word": "string",
                "start_word": "string",
                "summary": "string"
              }
            ],
            "topics": [
              {
                "end_word": "string",
                "start_word": "string",
                "text": "string",
                "topics": [
                  "string"
                ]
              }
            ],
            "transcript": "string",
            "words": [
              {
                "confidence": "string",
                "end": "string",
                "speaker": 0,
                "speaker_confidence": "string",
                "start": "string",
                "word": "string"
              }
            ]
          }
        ],
        "detected_language": "string",
        "search": [
          {
            "hits": [
              {
                "confidence": "string",
                "end": "string",
                "snippet": "string",
                "start": "string"
              }
            ],
            "query": "string"
          }
        ]
      }
    ],
    "intents": {
      "segments": [
        {
          "end_word": 0
        }
      ]
    }
  }
}
Documentation menu