Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Build a Voice Agent

Learn how to build a real-time voice agent using Deepgram’s Agent API.

Deepgram’s Voice Agent API uses a single WebSocket connection to handle the entire conversational loop. The API integrates speech-to-text, a large language model (LLM), and text-to-speech into one stream.

Building a voice agent involves four main steps over a WebSocket:

  1. Open a connection: Connect to the Deepgram Agent endpoint, wss://agent.deepgram.com/v1/agent/converse, using a supported SDK or a WebSocket client.
  2. Configure the agent: Send a Settings message to define the models, voices, and behavior.
  3. Stream audio: Send raw audio data to the agent.
  4. Handle events: Listen for transcripts, agent responses, and audio output.

Select a language to start building your voice agent. Each tutorial provides a complete, end-to-end implementation.

Once you understand the basics, you can explore more advanced configurations:

Check out these repositories for more complex voice agent implementations:

Use case Runtime / Language Repo
Basic demo Node, TypeScript, JavaScript Deepgram Voice Agent Demo
Medical assistant Node, TypeScript, JavaScript Medical Assistant Demo
Twilio integration Python Twilio Voice Agent (guide)
Text input demo Node, TypeScript, JavaScript Conversational AI Demo
Azure OpenAI Python Voice Agent with OpenAI Azure
Function calling Python / Flask Flask Agent Function Calling Demo

For information on concurrency limits, refer to the API Rate Limits documentation.

Deepgram calculates usage based on WebSocket connection time. One hour of connection time equals one hour of API usage.

Deepgram API Playground

Try this feature out in our API Playground.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu