Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Deploy Voice Agent

This guide covers deploying Deepgram’s Voice Agent API in a self-hosted Kubernetes environment using the Deepgram Helm chart. The Voice Agent API orchestrates Speech-to-Text (STT), a Large Language Model (LLM), and Text-to-Speech (TTS) into a single WebSocket-based conversational pipeline.

Voice Agent configuration files are available in the self-hosted-resources repository.

Two files are provided for AWS deployments:

  • 05-voice-agent-aws.cluster-config.yaml — EKS cluster configuration with dedicated node groups for API, Engine (GPU), and License Proxy workloads.
  • 05-voice-agent-aws.values.yaml — Helm values file with Voice Agent enabled, including STT, TTS, and end-of-turn Engine replicas.
Bash
BASE_URL="https://raw.githubusercontent.com/deepgram/self-hosted-resources/refs/heads/main"
curl -sSL "$BASE_URL/charts/deepgram-self-hosted/samples/05-voice-agent-aws.cluster-config.yaml" -o cluster-config.yaml
curl -sSL "$BASE_URL/charts/deepgram-self-hosted/samples/05-voice-agent-aws.values.yaml" -o values.yaml

Obtain your Voice Agent model links from your Deepgram Account Representative. The recommended models for Voice Agent deployments include:

  • If using Flux:

    • Flux model (e.g. flux-general-en.*.dg)
  • If using Nova-3:

    • Nova-3 model (e.g. nova-3-general.en.streaming.*.dg)
    • End-of-turn model (e.g. end-of-turn.*.dg)
  • If using Aura-2:

    • Voice model (e.g., aura-2.voice-pack.en.*.dg)
    • Generator model (e.g., aura-2.generator.en.*.dg)
  • If using Aura-1:

    • Voice model (e.g. aura-asteria-en.*.dg)
    • Phonemizer (e.g. phonemizer.en.*.dg)

You need a running Kubernetes cluster with GPU-enabled nodes. Choose a platform option based on your infrastructure and follow the corresponding setup guide:

Check that all pods are running:

Bash
kubectl get pods
# Confirm that api, engine (STT, TTS, EOT), and license-proxy pods are Running
kubectl logs <POD_NAME>

To make sure your Deepgram self-hosted Voice Agent deployment is properly configured and running, you will want to verify the services and make sample requests.

Forward the API service to test locally:

Bash
kubectl port-forward svc/deepgram-api 8080:8080

Unless you have HTTPS or TLS running on your API instance, construct your Deepgram API endpoint with http://, not https://, and ws://, not wss:// (for instance, ws://localhost:8080/v1/agent/converse).

Test the Speech-to-Text service with a sample audio file.

  1. Download a sample file from Deepgram (or supply your own file).

    Bash
    wget https://dpgr.am/bueller.wav
  2. Send your audio file to your local Deepgram setup for transcription.

    Bash
    curl -X POST --data-binary @bueller.wav "http://localhost:8080/v1/listen?model=nova-3"

You should receive a JSON response with the transcription and associated metadata.

Test the Text-to-Speech service with a sample speak request.

Bash
curl --request POST \
   --header "Content-Type: application/json" \
   --output tts-test.wav \
   --data '{"text":"This is a TTS self-hosted test."}' \
   --url "http://localhost:8080/v1/speak?model=aura-2-thalia-en"

You should receive a response with the audio output. You can play the file locally to evaluate the synthesized speech.

The Voice Agent Getting Started guide walks through building a voice agent with Deepgram’s hosted API. For self-hosted deployments, the only difference is how you initialize the DeepgramClient — you point it at your self-hosted endpoint instead of the default api.deepgram.com.

Hosted (default):

Python
api_key = os.getenv("DEEPGRAM_API_KEY")
client = DeepgramClient(api_key=api_key)

with client.agent.v1.connect() as connection:
    print("Created WebSocket connection...")

Self-hosted — set the baseUrl to your port-forwarded or load-balanced endpoint:

Python
import os

from deepgram import DeepgramClient
from deepgram.environment import DeepgramClientEnvironment

self_hosted_env = DeepgramClientEnvironment(
    base="http://localhost:8080",
    production="ws://localhost:8080",
    agent="ws://localhost:8080",
    agent_rest="http://localhost:8080"  # requires Python SDK v7.2.0+
)

api_key = os.getenv("DEEPGRAM_API_KEY")

client = DeepgramClient(
    api_key=api_key,
    environment=self_hosted_env
)

with client.agent.v1.connect() as connection:
    print("Created WebSocket connection...")

Once connected, follow the same steps as the Voice Agent Getting Started guide to configure and interact with your agent. For additional SDK configuration details, see Using SDKs with Self-Hosted.


What’s Next

Now that you have a Voice Agent deployment working, explore advanced configuration and integration options:

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu