Speech Started
vad_events boolean.
Pre-recorded Streaming:Nova All available languages
Deepgram’s Speech Started feature can be used for speech detection and can be used to detect the start of speech while transcribing live streaming audio.
SpeechStarted complements Voice Activity Detection (VAD) to promptly detect the start of speech post-silence. By gauging tonal nuances in human speech, the VAD can effectively differentiate between silent and non-silent audio segments, providing immediate notification of speech detection.
Enable Feature
Section titled “Enable Feature”To enable the SpeechStarted event, include the parameter vad_events=true in your request:
vad_events=true
You’ll then begin receiving messages upon speech starting.
# For more Python SDK migration guides, visit:
# https://github.com/deepgram/deepgram-python-sdk/tree/main/docs
with client.listen.v1.connect(
model="nova-3",
language="en-US",
# Apply smart formatting to the output
smart_format=True,
# Raw audio format details
encoding="linear16",
channels=1,
sample_rate=16000,
# To get UtteranceEnd, the following must be set:
interim_results=True,
utterance_end_ms="1000",
vad_events=True,
# Time in milliseconds of silence to wait for before finalizing speech
endpointing=300
) as connection:Results
Section titled “Results”The JSON message sent when the start of speech is detected looks similar to this:
{
"type": "SpeechStarted",
"channel": [
0,
1
],
"timestamp": 9.54
}- The
typefield is alwaysSpeechStartedfor this event. - The
channelfield is interpreted as[A,B], whereAis the channel index, andBis the total number of channels. The above example is channel 0 of single-channel audio. - The
timestampfield is the time at which speech was first detected.