Build a Voice Agent with LiveKit and Deepgram
If you already use LiveKit for WebRTC audio transport, you can add Deepgram’s speech-to-text and text-to-speech models to your LiveKit agent pipeline. LiveKit’s framework requires separate STT, LLM, and TTS providers, so this guide pairs Deepgram’s audio models with OpenAI for language understanding — though any LiveKit-compatible LLM works.
For a standalone voice agent without LiveKit or an external LLM, see the Deepgram Voice Agent API, which bundles STT, LLM routing, and TTS in a single WebSocket connection.
Before You Begin
Section titled “Before You Begin”This guide assumes you are familiar with Python or Node.js and have a basic understanding of how voice agents work.
Get OpenAI Credentials
Section titled “Get OpenAI Credentials”This tutorial uses OpenAI for its LLM. You’ll need to sign up for an OpenAI account and obtain an API key.
Get LiveKit Credentials
Section titled “Get LiveKit Credentials”You’ll need a LiveKit Cloud account with your LiveKit URL, API Key, and API Secret.
Requirements
Section titled “Requirements”- Python 3.10+ or Node.js 18+
Set Up Your Project
Section titled “Set Up Your Project”This implementation is a starting reference for building your own voice agent with LiveKit and Deepgram. It is not designed for production deployments.
Create a new directory, set up your environment, and install the LiveKit agents framework along with the Deepgram and Silero plugins:
mkdir deepgram-livekit-agent
cd deepgram-livekit-agent
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
pip install "livekit-agents[openai]" livekit-plugins-deepgram livekit-plugins-silero python-dotenvSet Environment Variables
Section titled “Set Environment Variables”Create a .env file in your project root with the credentials you collected earlier. The agent reads these at startup to authenticate with each service:
DEEPGRAM_API_KEY=your_deepgram_api_key
OPENAI_API_KEY=your_openai_api_key
LIVEKIT_URL=wss://your-project.livekit.cloud
LIVEKIT_API_KEY=your_livekit_api_key
LIVEKIT_API_SECRET=your_livekit_api_secretLIVEKIT_URL is the WebSocket endpoint for your LiveKit Cloud project. You can find it along with your API key and secret in the LiveKit dashboard.
Build the Agent
Section titled “Build the Agent”The agent connects to a LiveKit room, creates a session with Deepgram for audio processing and OpenAI for language understanding, and starts listening for speech.
The key components are:
Agent— defines the agent’s personality and instructions.AgentSession— wires together the STT, LLM, TTS, and VAD providers into a pipeline. When a user speaks, audio flows through Deepgram Nova-3 for transcription, OpenAI GPT-4o for a response, and Deepgram Aura for speech synthesis.generate_reply— triggers the agent’s first message so it greets the user without waiting for input.
Create agent.py (Python) or agent.ts (Node.js):
# agent.py
from dotenv import load_dotenv
from livekit.agents import (
Agent,
AgentSession,
AgentServer,
JobContext,
cli,
)
from livekit.plugins import deepgram, openai, silero
load_dotenv()
server = AgentServer()
class VoiceAssistant(Agent):
def __init__(self):
super().__init__(
instructions=(
"You are a friendly, helpful voice assistant. "
"Keep your responses concise — aim for 1-3 sentences "
"unless the user asks for detail."
),
)
@server.rtc_session()
async def entrypoint(ctx: JobContext):
await ctx.connect()
session = AgentSession(
stt=deepgram.STT(
model="nova-3",
language="en",
punctuate=True,
smart_format=True,
interim_results=True,
),
llm=openai.LLM(model="gpt-4o"),
tts=deepgram.TTS(model="aura-2-thalia-en"),
vad=silero.VAD.load(),
)
await session.start(
agent=VoiceAssistant(),
room=ctx.room,
)
await session.generate_reply(
instructions="Greet the user and ask how you can help.",
allow_interruptions=True,
)
if __name__ == "__main__":
cli.run_app(server)Run the Agent
Section titled “Run the Agent”Start the agent in development mode. The dev flag connects the agent to your LiveKit Cloud project and automatically registers it to handle incoming sessions:
python agent.py devTest the Agent
Section titled “Test the Agent”LiveKit provides the Agents Playground — a browser-based tool for testing agents. It includes video, chat, and other features, but for this tutorial you only need the microphone.
- Go to agents-playground.livekit.io
- Enter your LiveKit Cloud URL and a participant token, then connect
- Allow microphone access when prompted
- Start talking — the agent should respond in real time
You can generate a participant token from the LiveKit dashboard or using the LiveKit CLI.
The agent greets you automatically on connect. Silero VAD detects when you stop speaking and triggers the STT-to-LLM-to-TTS pipeline. You can interrupt the agent mid-sentence — VAD handles barge-in automatically.
Further Reading
Section titled “Further Reading”- Deepgram Voice Agent API — build voice agents without an external LLM or transport layer
- LiveKit Agents Documentation — LiveKit’s agent framework reference