Getting Started
Deepgram API Playground
Try this feature out in our API Playground.
This guide walks you through transcribing pre-recorded audio with the Deepgram API using cURL or one of Deepgram’s SDKs.
Replace YOUR_DEEPGRAM_API_KEY with your API key and run the following in a terminal or API client.
Remote file
Section titled “Remote file”curl \
--request POST \
--header 'Authorization: Token YOUR_DEEPGRAM_API_KEY' \
--header 'Content-Type: application/json' \
--data '{"url":"https://dpgr.am/spacewalk.wav"}' \
--url 'https://api.deepgram.com/v1/listen?model=nova-3&smart_format=true'Local file
Section titled “Local file”Replace @youraudio.wav with the path to an audio file on your computer. See Supported Audio Formats for accepted formats.
curl \
--request POST \
--header 'Authorization: Token YOUR_DEEPGRAM_API_KEY' \
--header 'Content-Type: audio/wav' \
--data-binary @youraudio.wav \
--url 'https://api.deepgram.com/v1/listen?model=nova-3&smart_format=true'To transcribe pre-recorded audio using one of Deepgram’s SDKs, follow these steps.
Install the SDK and dependencies
Section titled “Install the SDK and dependencies”Open your terminal, navigate to your project directory, and install the Deepgram SDK along with any required dependencies.
# Install the Deepgram JS SDK and dotenv
# https://github.com/deepgram/deepgram-js-sdk
npm install @deepgram/sdk dotenvTranscribe a remote file
Section titled “Transcribe a remote file”Create a new file in your project and add the following code to transcribe a remote audio file by URL:
// index.js (node example)
const { DeepgramClient } = require("@deepgram/sdk");
const dotenv = require("dotenv");
dotenv.config();
const transcribeUrl = async () => {
// STEP 1: Create a Deepgram client using the API key
const deepgram = new DeepgramClient({ apiKey: process.env.DEEPGRAM_API_KEY });
// STEP 2: Call the transcribeUrl method with the audio payload and options
// STEP 3: Configure Deepgram options for audio analysis
const result = await deepgram.listen.v1.media.transcribeUrl({
url: "https://dpgr.am/spacewalk.wav",
model: "nova-3",
smart_format: true,
});
// STEP 4: Print the results
console.dir(result, { depth: null });
};
transcribeUrl();Transcribe a local file
Section titled “Transcribe a local file”// index.js (node example)
const { DeepgramClient } = require("@deepgram/sdk");
const fs = require("fs");
const dotenv = require("dotenv");
dotenv.config();
const transcribeFile = async () => {
// STEP 1: Create a Deepgram client using the API key
const deepgram = new DeepgramClient({ apiKey: process.env.DEEPGRAM_API_KEY });
// STEP 2: Call the transcribeFile method with the audio payload and options
// STEP 3: Configure Deepgram options for audio analysis
const result = await deepgram.listen.v1.media.transcribeFile(
// path to the audio file
fs.createReadStream("spacewalk.mp3"),
{
model: "nova-3",
smart_format: true,
}
);
// STEP 4: Print the results
console.dir(result, { depth: null });
};
transcribeFile();Non-SDK code examples
Section titled “Non-SDK code examples”For language-specific examples that don’t use Deepgram’s SDKs, see the recipes repository. We recommend trying the SDKs first.
Results
Section titled “Results”Run your application from the terminal. Your transcript appears in your shell.
node index.jsAnalyze the response
Section titled “Analyze the response”When the file finishes processing (often after only a few seconds), you receive a JSON response:
{
"metadata": {
"transaction_key": "deprecated",
"request_id": "2479c8c8-8185-40ac-9ac6-f0874419f793",
"sha256": "154e291ecfa8be6ab8343560bcc109008fa7853eb5372533e8efdefc9b504c33",
"created": "2024-02-06T19:56:16.180Z",
"duration": 25.933313,
"channels": 1,
"models": [
"30089e05-99d1-4376-b32e-c263170674af"
],
"model_info": {
"30089e05-99d1-4376-b32e-c263170674af": {
"name": "2-general-nova",
"version": "2024-01-09.29447",
"arch": "nova-3"
}
}
},
"results": {
"channels": [
{
"alternatives": [
{
"transcript": "Yeah. As as much as, it's worth celebrating, the first, spacewalk, with an all female team, I think many of us are looking forward to it just being normal. And, I think if it signifies anything, It is, to honor the the women who came before us who, were skilled and qualified, and didn't get the the same opportunities that we have today.",
"confidence": 0.99902344,
"words": [
{
"word": "yeah",
"start": 0.08,
"end": 0.32,
"confidence": 0.9975586,
"punctuated_word": "Yeah."
},
{
"word": "as",
"start": 0.32,
"end": 0.79999995,
"confidence": 0.9921875,
"punctuated_word": "As"
}
],
"paragraphs": {
"transcript": "\nYeah. As as much as, it's worth celebrating...",
"paragraphs": [
{
"sentences": [
{
"text": "Yeah.",
"start": 0.08,
"end": 0.32
}
],
"num_words": 63,
"start": 0.08,
"end": 25.52
}
]
}
}
]
}
]
}
}In this response:
-
transcript: the transcript for the audio segment being processed. -
confidence: a floating point value between 0 and 1 that indicates overall transcript reliability. Larger values indicate higher confidence. -
words: an object containing eachwordin the transcript, along with itsstarttime andendtime (in seconds) from the beginning of the audio stream, and aconfidencevalue.- Because we passed the
smart_format: trueoption, each word object also includes itspunctuated_wordvalue, which contains the transformed word after punctuation and capitalization are applied.
- Because we passed the
Limits
Section titled “Limits”- File size: Maximum 2 GB. For large video files, extract the audio stream first.
- Rate limits: Up to 100 concurrent requests per project for Nova, Base, and Enhanced models. For full details, see API Rate Limits.
- Processing time: Requests exceeding 10 minutes (Nova/Base/Enhanced) or 20 minutes (Whisper) return a
504: Gateway Timeouterror.
What’s next?
Section titled “What’s next?”- Feature overview: Review the full list of features available for pre-recorded speech-to-text.
- Language: Transcribe audio in other languages.
- Streaming audio: Transcribe audio in real time.
- Use cases: Explore ways to use Deepgram products.