Getting Started
Deepgram API Playground
Try this feature out in our API Playground.
In this guide, you’ll learn how to automatically transcribe live streaming audio in real time using Deepgram’s SDKs, which are supported for use with the Deepgram API. (If you prefer not to use a Deepgram SDK, jump to the section Non-SDK Code Examples.)
To transcribe audio from an audio stream using one of Deepgram’s SDKs, follow these steps.
Install the SDK
Section titled “Install the SDK”Open your terminal, navigate to the location on your drive where you want to create your project, and install the Deepgram SDK.
// Install the Deepgram JS SDK
// https://github.com/deepgram/deepgram-js-sdk
// npm install @deepgram/sdkAdd Dependencies
Section titled “Add Dependencies”// Install cross-fetch: Platform-agnostic Fetch API with typescript support, a simple interface, and optional polyfill.
// Install dotenv to protect your api key
// $ npm install cross-fetch dotenvTranscribe Audio from a Remote Stream
Section titled “Transcribe Audio from a Remote Stream”The following code shows how to transcribe audio from a remote audio stream.
// Example filename: index.js
const { DeepgramClient } = require("@deepgram/sdk");
const fetch = require("cross-fetch");
const dotenv = require("dotenv");
dotenv.config();
// URL for the realtime streaming audio you would like to transcribe
const url = "http://stream.live.vc.bbcmedia.co.uk/bbc_world_service";
const live = async () => {
// STEP 1: Create a Deepgram client using the API key
const deepgram = new DeepgramClient({ apiKey: process.env.DEEPGRAM_API_KEY });
// STEP 2: Create a live transcription connection
const connection = await deepgram.listen.v1.connect({
model: "nova-3",
language: "en-US",
smart_format: "true",
});
// STEP 3: Listen for events from the live transcription connection
connection.on("open", () => {
connection.on("close", () => {
console.log("Connection closed.");
});
connection.on("message", (data) => {
if (data.type === "Results") {
console.log(data.channel.alternatives[0].transcript);
}
});
connection.on("error", (err) => {
console.error(err);
});
// STEP 4: Fetch the audio stream and send it to the live transcription connection
fetch(url)
.then((r) => r.body)
.then((res) => {
res.on("readable", () => {
connection.sendMedia(res.read());
});
});
});
connection.connect();
await connection.waitForOpen();
};
live();Non-SDK Code Examples
Section titled “Non-SDK Code Examples”If you would like to try out making a Deepgram speech-to-text request in a specific language (but not using Deepgram’s SDKs), we offer a library of code-samples in this Github repo. However, we recommend first trying out our SDKs.
Results
Section titled “Results”In order to see the results from Deepgram, you must run the application. Run your application from the terminal. Your transcripts will appear in your shell.
# Run your application using the file you created in the previous step
# Example: node index.js
node YOUR_FILE_NAME.jsAnalyze the Response
Section titled “Analyze the Response”The responses that are returned will look similar to this:
{
"type": "Results",
"channel_index": [
0,
1
],
"duration": 1.98,
"start": 5.99,
"is_final": true,
"speech_final": true,
"channel": {
"alternatives": [
{
"transcript": "Tell me more about this.",
"confidence": 0.99964225,
"words": [
{
"word": "tell",
"start": 6.0699997,
"end": 6.3499994,
"confidence": 0.99782443,
"punctuated_word": "Tell"
},
{
"word": "me",
"start": 6.3499994,
"end": 6.6299996,
"confidence": 0.9998324,
"punctuated_word": "me"
},
{
"word": "more",
"start": 6.6299996,
"end": 6.79,
"confidence": 0.9995466,
"punctuated_word": "more"
},
{
"word": "about",
"start": 6.79,
"end": 7.0299997,
"confidence": 0.99984455,
"punctuated_word": "about"
},
{
"word": "this",
"start": 7.0299997,
"end": 7.2699995,
"confidence": 0.99964225,
"punctuated_word": "this"
}
]
}
]
},
"metadata": {
"request_id": "52cc0efe-fa77-4aa7-b79c-0dda09de2f14",
"model_info": {
"name": "2-general-nova",
"version": "2024-01-18.26916",
"arch": "nova-2"
},
"model_uuid": "c0d1a568-ce81-4fea-97e7-bd45cb1fdf3c"
},
"from_finalize": false
}In this default response, we see:
-
transcript: the transcript for the audio segment being processed. -
confidence: a floating point value between 0 and 1 that indicates overall transcript reliability. Larger values indicate higher confidence. -
words: an object containing eachwordin the transcript, along with itsstarttime andendtime (in seconds) from the beginning of the audio stream, and aconfidencevalue.- Because we passed the
smart_format: trueoption to thetranscription.prerecordedmethod, each word object also includes itspunctuated_wordvalue, which contains the transformed word after punctuation and capitalization are applied.
- Because we passed the
-
speech_final: tells us this segment of speech naturally ended at this point. By default, Deepgram live streaming looks for any deviation in the natural flow of speech and returns a finalized response at these places. To learn more about this feature, see Endpointing. -
is_final: If this saysfalse, it is indicating that Deepgram will continue waiting to see if more data will improve its predictions. Deepgram live streaming can return a series of interim transcripts followed by a final transcript. To learn more, see Interim Results.
If your scenario requires you to keep the connection alive even while data is not being sent to Deepgram, you can send periodic KeepAlive messages to essentially “pause” the connection without closing it. To learn more, see KeepAlive.
What’s Next?
Section titled “What’s Next?”Now that you’ve gotten transcripts for streaming audio, enhance your knowledge by exploring the following areas. You can also check out our Live Streaming API Reference for a list of all possible parameters.
Read the Feature Guides
Section titled “Read the Feature Guides”Deepgram’s features help you to customize your transcripts.
- Language: Learn how to transcribe audio in other languages.
- Feature Overview: Review the list of features available for streaming speech-to-text. Then, dive into individual guides for more details.
Tips and tricks
Section titled “Tips and tricks”- End of speech detection - Learn how to pinpoint end of speech post-speaking more effectively.
- Using interim results - Learn how to use preliminary results provided during the streaming process which can help with speech detection.
- Measuring streaming latency - Learn how to measure latency in real-time streaming of audio.
Add Your Audio
Section titled “Add Your Audio”- Ready to connect Deepgram to your own audio source? Start by reviewing how to determine your audio format and format your API request accordingly.
- Then, check out our Live Streaming Starter Kit. It’s the perfect “102” introduction to integrating your own audio.
Explore Use Cases
Section titled “Explore Use Cases”- Learn about the different ways you can use Deepgram products to help you meet your business objectives. Explore Deepgram’s use cases.
Transcribe Pre-recorded Audio
Section titled “Transcribe Pre-recorded Audio”- Now that you know how to transcribe streaming audio, check out how you can use Deepgram to transcribe pre-recorded audio. To learn more, see Getting Started with Pre-recorded Audio.