Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Getting Started

This guide will walk you through how to turn streaming text into speech with Deepgram’s text-to-speech Websocket API.

Deepgram has several SDKs that can make the API easier to use. Follow these steps to use the SDK of your choice to make a Deepgram TTS request.

Bash
# Install the SDK
npm install @deepgram/sdk

# Add the dependencies
npm install dotenv
JavaScript
const fs = require("fs");
const { DeepgramClient } = require("@deepgram/sdk");

// Add a wav audio container header to the file if you want to play the audio
// using the AudioContext or media player like VLC, Media Player, or Apple Music
// Without this header in the Chrome browser case, the audio will not play.
// prettier-ignore
const wavHeader = [
0x52, 0x49, 0x46, 0x46, // "RIFF"
0x00, 0x00, 0x00, 0x00, // Placeholder for file size
0x57, 0x41, 0x56, 0x45, // "WAVE"
0x66, 0x6D, 0x74, 0x20, // "fmt "
0x10, 0x00, 0x00, 0x00, // Chunk size (16)
0x01, 0x00,             // Audio format (1 for PCM)
0x01, 0x00,             // Number of channels (1)
0x80, 0xBB, 0x00, 0x00, // Sample rate (48000)
0x00, 0xEE, 0x02, 0x00, // Byte rate (48000 * 2)
0x02, 0x00,             // Block align (2)
0x10, 0x00,             // Bits per sample (16)
0x64, 0x61, 0x74, 0x61, // "data"
0x00, 0x00, 0x00, 0x00  // Placeholder for data size
];

const live = async () => {
const text = "Hello, how can I help you today?";

const deepgram = new DeepgramClient({ apiKey: process.env.DEEPGRAM_API_KEY });

const dgConnection = await deepgram.speak.v1.connect({
  model: "aura-2-thalia-en",
  encoding: "linear16",
  sample_rate: 48000,
});

let audioBuffer = Buffer.from(wavHeader);

dgConnection.on("open", () => {
  console.log("Connection opened");

  // Send text data for TTS synthesis
  dgConnection.sendText({ type: "Text", text });

  dgConnection.on("close", () => {
    console.log("Connection closed");
  });

  dgConnection.on("message", (data) => {
    if (data.type === "Metadata") {
      console.dir(data, { depth: null });
    } else if (data.type === "Flushed") {
      console.log("Deepgram Flushed");
      // Write the buffered audio data to a file when the flush event is received
      writeFile();
    } else if (typeof data === "string") {
      console.log("Deepgram audio data received");
      // Concatenate the audio chunks into a single buffer
      const buffer = Buffer.from(data, "base64");
      audioBuffer = Buffer.concat([audioBuffer, buffer]);
    }
  });

  dgConnection.on("error", (err) => {
    console.error(err);
  });
});

dgConnection.connect();
await dgConnection.waitForOpen();

const writeFile = () => {
  if (audioBuffer.length > 0) {
    fs.writeFile("output.wav", audioBuffer, (err) => {
      if (err) {
        console.error("Error writing audio file:", err);
      } else {
        console.log("Audio file saved as output.wav");
      }
    });
    audioBuffer = Buffer.from(wavHeader); // Reset buffer after writing
  }
};
};

live();

Below is a high-level workflow for obtaining an audio stream from user-provided text.

To establish a connection, you must provide a few parameters on the URL to describe the type of audio you want. You can check out the API Reference to set the audio model, which controls the voice, the encoding, and the sample rate of the audio.

Send the desired text to transform to audio using the WebSocket message below:

JSON
{
  "type": "Speak",
  "text": "Your text to transform to speech",
}

When you have queued enough text, you can obtain the corresponding audio by sending a Flush command.

JSON
{
  "type": "Flush"
}

Upon successfully sending the Flush, you will receive an audio byte stream from the websocket connection containing the synthesized text-to-speech. The format will be based on the encoding values provided upon establishing the connection.

When you are finished with the WebSocket, you can close the connection by sending the following Close command.

JSON
{
  "type": "Close"
}

Keep these limits in mind when making a Deepgram text-to-speech request.

If you are building for conversational AI use cases where a human is talking to a TTS agent, a single websocket per conversation is required. After you establish a connection, you will not be able to change the voice or media output settings.

Sending a request with a text payload longer than the maximum number of characters can result in a 413: Input Text Exceeds Character Limits error, and the audio file will not be created.

Model Max Characters
Aura-2, Aura-1 2000

A 422: Unprocessable Content error can be returned if the client fails to send the request successfully.

The throughput limit is 2400 characters per minute and is measured by the number of characters sent to the websocket.

An active websocket has a 60-minute timeout period from the initial connection. This timeout exists for connections that are actively being used. If you desire a connection for longer than 60 minutes, create a new websocket connection to Deepgram.

The maximum number of times you can send the Flush message is 20 times every 60 seconds. After that, you will receive a warning message stating that we cannot process any more flush messages until the 60-second time window has passed.

If the number of in-progress requests for a project meets or exceeds the rate limit, new requests will receive a 429: Too Many Requests error.

Now that you’ve transformed text into speech with Deepgram’s API, enhance your knowledge by exploring the following areas.

Deepgram’s features help you customize your request to produce the best output for your use case. Here are a few guides that can help:

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu