Migrating from AssemblyAI Speech-to-Text to Deepgram
This guide provides a detailed step-by-step process for developers transitioning from AssemblyAI speech-to-text (STT) services to Deepgram’s STT services using the Deepgram SDKs. The goal is to ensure a smooth migration by highlighting differences and demonstrating equivalent functionalities between the two platforms.
Getting Started
Section titled “Getting Started”Prerequisites
Section titled “Prerequisites”Before proceeding with the migration, ensure you meet the following prerequisites:
Required Tools
Section titled “Required Tools”- A code editor (e.g., Visual Studio Code)
- Terminal or command prompt access
- Node / Python installed
API Keys
Section titled “API Keys”- AssemblyAI API Key: Obtain from your AssemblyAI dashboard.
- Deepgram API Key: Sign up on Deepgram’s platform and get your API key in the Deepgram Console.
Overview of AssemblyAI and Deepgram APIs
Section titled “Overview of AssemblyAI and Deepgram APIs”Both AssemblyAI and Deepgram provide robust speech-to-text APIs, but they have different endpoints, request parameters, and response structures. This guide will map AssemblyAI functionalities to their Deepgram equivalents.
Step-by-Step Migration Instructions
Section titled “Step-by-Step Migration Instructions”\1. Setting Up the Environment
Section titled “\1. Setting Up the Environment”Node: Ensure you have Node.js installed. If not, download and install it from the Node.js website.
Python: Ensure you have Python installed. If not, download and install it from the Python website.
\2. Configuring API Keys
Section titled “\2. Configuring API Keys”How to configure AssemblyAI API key
Section titled “How to configure AssemblyAI API key”Create a .env file in your project directory and add your AssemblyAI API key:
ASSEMBLYAI_API_KEY=your_assemblyai_api_key_hereHow to configure Deepgram API key
Section titled “How to configure Deepgram API key”Similarly, add your Deepgram API key to the .env file:
DEEPGRAM_API_KEY=your_deepgram_api_key_here\3. Installing the SDK and Dependencies
Section titled “\3. Installing the SDK and Dependencies”For AssemblyAI:
npm install assemblyai dotenvFor Deepgram:
npm install @deepgram/sdk dotenv\4. Making API Requests
Section titled “\4. Making API Requests”Initialization
Section titled “Initialization”AssemblyAI Initialization:
import { AssemblyAI } from "assemblyai";
import dotenv from "dotenv";
dotenv.config();
const client = new AssemblyAI({
apiKey: process.env.ASSEMBLYAI_API_KEY,
});Deepgram Initialization:
import { DeepgramClient } from "@deepgram/sdk";
import dotenv from "dotenv";
dotenv.config();
const deepgram = new DeepgramClient({ apiKey: process.env.DEEPGRAM_API_KEY });Add Request Parameters
Section titled “Add Request Parameters”AssemblyAI:
const data = {
audio_url: "https://dpgr.am/spacewalk.wav", // the audio_url for the audio being transcribed is included
speech_model: "nano",
speaker_labels: true,
};Deepgram:
const options = {
model: "nova-3",
smart_format: true,
// Do not include the audio_url in this object
};Example: Transcribe Audio Using a Remote URL
Section titled “Example: Transcribe Audio Using a Remote URL”Here is the entire code sample that shows how to transcribe audio using a remote URL.
AssemblyAI:
import { AssemblyAI } from "assemblyai";
import dotenv from "dotenv";
dotenv.config();
const client = new AssemblyAI({
apiKey: process.env.ASSEMBLYAI_API_KEY,
});
const FILE_URL = "https://dpgr.am/spacewalk.wav";
const data = {
audio_url: FILE_URL,
speech_model: "nano",
speaker_labels: true,
};
const run = async () => {
const response = await client.transcripts.transcribe(data);
console.log(JSON.stringify(response));
};
run();Deepgram:
import { DeepgramClient } from "@deepgram/sdk";
import dotenv from "dotenv";
dotenv.config();
const run = async () => {
const deepgram = new DeepgramClient({ apiKey: process.env.DEEPGRAM_API_KEY });
const response = await deepgram.listen.v1.media.transcribeUrl({
url: "https://dpgr.am/spacewalk.wav",
model: "nova-3",
diarize: true,
});
console.dir(JSON.stringify(response), { depth: null });
};
run();Example: Transcribe Audio Using a Local File
Section titled “Example: Transcribe Audio Using a Local File”Here is the entire code sample that shows how to transcribe audio using a local file.
AssemblyAI:
import { AssemblyAI } from "assemblyai";
import dotenv from "dotenv";
dotenv.config();
const client = new AssemblyAI({
apiKey: process.env.ASSEMBLYAI_API_KEY,
});
const AUDIO_FILE = "sample.wav";
const data = {
audio: AUDIO_FILE,
speech_model: "nano",
speaker_labels: true,
};
const run = async () => {
const response = await client.transcripts.transcribe(data);
console.log(JSON.stringify(response));
};
run();Deepgram:
import { DeepgramClient } from "@deepgram/sdk";
import fs from "fs";
import dotenv from "dotenv";
dotenv.config();
const deepgram = new DeepgramClient({ apiKey: process.env.DEEPGRAM_API_KEY });
const run = async () => {
const response = await deepgram.listen.v1.media.transcribeFile(
fs.createReadStream("sample.wav"),
{
model: "nova-3",
diarize: true,
}
);
console.dir(JSON.stringify(response), { depth: null });
};
run();\7. Handling Responses
Section titled “\7. Handling Responses”Compare the JSON responses:
Section titled “Compare the JSON responses:”AssemblyAI:
{
"id": "some_id",
"status": "completed",
"audio_url": "https://dpgr.am/spacewalk.wav",
"text": "Transcript text here...",
"words": [
{
"start": 255,
"end": 767,
"text": "Yeah.",
"confidence": 0.97465,
"speaker": null
},
]
}Deepgram:
{
"metadata": {
"transaction_key": "deprecated",
"request_id": "unique_request_id",
"created": "2024-02-06T19:56:16.180Z",
"duration": 25.933313,
"channels": 1,
"models": ["1abfe86b-e047-4eed-858a-35e5625b41ee"],
"model_info": {}
},
"results": {
"channels": [
{
"alternatives": [
{
"transcript": "Transcript text here...",
"confidence": 0.99902344,
"words": [
{
"word": "yeah",
"start": 0.08,
"end": 0.32,
"confidence": 0.9975586,
"punctuated_word": "Yeah."
}
]
}
]
}
]
}
}\8. Code Migration
Section titled “\8. Code Migration”Adapting code to handle Deepgram’s response structure involves accessing nested fields within the JSON response. For instance, response.results.channels[0].alternatives[0].transcript will give you the transcript text.
Be sure to update your data parsing logic to correctly navigate the nested response format, and thoroughly test the new code to ensure it handles various edge cases and accurately extracts the needed information.
\9. Testing and Validation
Section titled “\9. Testing and Validation”Steps to test the integration
Section titled “Steps to test the integration”- Run the application to generate transcriptions.
- Validate that the responses match expected outputs.
Validating transcription accuracy and performance
Section titled “Validating transcription accuracy and performance”- Compare the transcript text from both AssemblyAI and Deepgram.
- Evaluate the confidence scores and accuracy of the transcriptions.