Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Voice Agent Message Flow

This guide walks you through implementing the correct message flow when building a Voice Agent client. Follow these steps to establish a connection, configure settings, and handle the conversation loop.

Establish the Connection and Receive Welcome

Section titled “Establish the Connection and Receive Welcome”
  1. Open a WebSocket connection to the Voice Agent endpoint.

  2. Wait for the server to send a Welcome message confirming the connection:

JSON
{ "type": "Welcome", "request_id": "uuid" }

Configure Settings and Wait for Confirmation

Section titled “Configure Settings and Wait for Confirmation”
  1. Send a Settings message with your audio and agent configuration:
JSON
{
  "type": "Settings",
  "audio": {
    "input": {
      "encoding": "linear16",
      "sample_rate": 16000
    },
    "output": {
      "encoding": "linear16",
      "sample_rate": 24000,
      "container": "none"
    }
  },
  "agent": {
    "listen": { "provider": { "type": "deepgram", "model": "nova-3" } },
    "think": {
      "provider": { "type": "open_ai", "model": "gpt-4o-mini" }
    },
    "speak": { "provider": { "type": "deepgram", "version": "v2", "model": "flux-kit-en" } }
  }
}
  1. Wait for the server to send a SettingsApplied message:
JSON
{ "type": "SettingsApplied" }
  1. After receiving SettingsApplied, begin streaming binary audio data (PCM) continuously to the server.

  2. Optionally, send text input using InjectUserMessage:

JSON
{ "type": "InjectUserMessage", "content": "Hello" }
  1. Process the following events as the conversation progresses:
Event Description
UserStartedSpeaking User began talking. Stop any audio playback immediately to handle barge-in.
ConversationText User's speech has been transcribed.
AgentThinking Agent is processing the user's input.
ConversationText Agent's text response is available.
[binary audio] Agent's audio response. Play this through your audio output.
AgentAudioDone Agent finished speaking.
Error / Warning Issues occurred during processing.

Confirm your implementation works correctly by checking:

  • You receive a Welcome message immediately after connecting.
  • You receive a SettingsApplied message after sending your Settings.
  • The agent responds with ConversationText and binary audio when you speak or inject text.
  • Audio playback stops when you receive UserStartedSpeaking (barge-in detection).
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu