LiveKit and Deepgram
This guide walks you through building a voice AI agent that uses LiveKit Agents for real-time audio transport and Deepgram for speech-to-text (STT) and text-to-speech (TTS). By the end, you will have a working voice agent that listens to a user, generates a response with an LLM, and speaks back in real time.
Deepgram is available in LiveKit Agents through two paths:
| Path | Description |
|---|---|
| LiveKit Inference | Deepgram models hosted and billed through LiveKit Cloud. Supports Nova-3, Nova-2 variants (nova-2, nova-2-medical, nova-2-phonecall), Flux, and Aura-2. No Deepgram API key required. |
| Deepgram Plugin | Connects directly to Deepgram’s API with your own API key. Includes diarization, keyterm prompting, and advanced parameters. |
This guide starts with LiveKit Inference for the fastest setup, then shows how to switch to the Deepgram Plugin for direct API access and advanced features.
Code samples are provided for both Python and TypeScript (Node.js). Choose one language and follow its tab in each step.
Before you begin
Section titled “Before you begin”Before you can use Deepgram, you need to create a Deepgram account. Signup is free and includes $200 in credit.
You need:
- A LiveKit Cloud account (or a self-hosted LiveKit server)
- An LLM. The starter uses LiveKit Inference (
inference.LLM(model="openai/chat-latest")), which runs through LiveKit Cloud and requires no separate provider key. A dedicated LLM API key is only needed if you bring your own provider, such as OpenAI; LiveKit Agents supports other providers too. - Python 3.10+ or Node.js 20+
Step 1: Create the project
Section titled “Step 1: Create the project”The fastest way to scaffold a new agent project is with the LiveKit CLI (lk). Alternatively, clone a starter template from GitHub (Python, Node.js).
When the CLI finishes, your agent is registered with LiveKit Cloud. You’ll test it later from the Agent Console.
Step 2: Configure your environment
Section titled “Step 2: Configure your environment”The CLI creates a .env.local file during setup. Open it and confirm your API keys are set:
If you used the CLI’s guided setup, these values are already populated. If not, add them from your LiveKit Cloud dashboard. The OPENAI_API_KEY is optional — the starter uses LiveKit Inference by default, so you only need it if you bring your own OpenAI key from the OpenAI dashboard.
Step 3: Use Deepgram for TTS
Section titled “Step 3: Use Deepgram for TTS”At the time of writing, the starter template uses Deepgram for STT but Cartesia for TTS. These defaults change frequently and aren’t guaranteed, so check the generated code. To use Deepgram for both, find the tts argument in the AgentSession constructor in src/agent.py (Python) or src/main.ts (Node.js) and replace it:
# Python — src/agent.py
# Replace the existing tts line with:
tts=inference.TTS(model="deepgram/aura-2", voice="thalia"),// TypeScript — src/main.ts
// Replace the existing tts line with:
tts: new inference.TTS({
model: 'deepgram/aura-2',
voice: 'thalia',
}),The agent now uses Deepgram for both STT and TTS. No additional dependencies or API keys are needed because both run through LiveKit Inference.
Browse available voices in the Deepgram voice library.
Step 4: Run the agent
Section titled “Step 4: Run the agent”Start the agent in development mode. The required model files (VAD, turn detection) are now downloaded automatically:
The dev command connects your agent to LiveKit Cloud. Open the Agent Console in your browser to talk to your agent.
Step 5: Test the conversation
Section titled “Step 5: Test the conversation”Once the agent is running, you should hear a greeting. Try these interactions to verify everything works:
- Ask a question and confirm the agent responds with speech.
- Start speaking while the agent is talking. It should stop and listen.
- Pause after speaking. The agent should detect the end of your turn and respond.
If any of these fail, check your API keys and confirm the agent process is running.
Use the Deepgram Plugin for advanced features
Section titled “Use the Deepgram Plugin for advanced features”The steps above use LiveKit Inference, which hosts Deepgram models through LiveKit Cloud. If you need direct access to Deepgram features like speaker diarization, keyterm prompting, or fine-grained parameter control, use the Deepgram Plugin instead. This connects directly to Deepgram’s API with your own API key.
Install the plugin
Section titled “Install the plugin”Set your Deepgram API key
Section titled “Set your Deepgram API key”Add DEEPGRAM_API_KEY to your .env.local file:
Replace YOUR_DEEPGRAM_API_KEY with the API key from your Deepgram Console. The plugin reads this variable automatically at startup.
Update the AgentSession
Section titled “Update the AgentSession”Add the Deepgram import at the top of your entrypoint file, then replace the stt and tts arguments in the AgentSession:
# Python — src/agent.py
from livekit.plugins import deepgram
# Replace the stt and tts lines in your AgentSession:
stt=deepgram.STT(
model="nova-3",
language="en",
punctuate=True,
interim_results=True,
),
tts=deepgram.TTS(model="aura-2-thalia-en"),// TypeScript — src/main.ts
import * as deepgram from "@livekit/agents-plugin-deepgram";
// Replace the stt and tts lines in your AgentSession:
stt: new deepgram.STT({
model: "nova-3",
language: "en",
punctuate: true,
interimResults: true,
}),
tts: new deepgram.TTS({ model: "aura-2-thalia-en" }),Your agent now connects directly to Deepgram’s API with your own API key, instead of going through LiveKit Inference. This unlocks the advanced features covered below, such as keyterm prompting, speaker diarization, and fine-grained STT/TTS parameters.
To verify the change, restart the agent as in Step 4 (uv run src/agent.py dev for Python, or pnpm run dev for Node.js) and run through the checks in Step 5. The conversation should behave as before, now powered by the Deepgram plugin. If the agent fails to start, confirm DEEPGRAM_API_KEY is set in .env.local.
Use Flux for turn detection
Section titled “Use Flux for turn detection”Flux is Deepgram’s conversational STT model with built-in turn detection. It uses acoustic and semantic cues to determine when a speaker has finished their turn, resulting in more natural conversations with fewer awkward pauses.
To use Flux, replace the stt configuration with STTv2 and set turn detection to "stt":
# Replace the stt line and add turn_handling:
stt=deepgram.STTv2(model="flux-general-en"),
turn_handling={"turn_detection": "stt"},// TypeScript — src/main.ts
// Replace the stt line and add turnDetection:
stt: new deepgram.STTv2({ model: "flux-general-en" }),
turnHandling: { turnDetection: "stt" },Choose a different voice
Section titled “Choose a different voice”Flux is the latest conversation-native TTS model, built for real-time voice agents. It’s expressive by default, consistent across turns, and responds in under 200ms with native interruption handling, real-time controls, and strong entity accuracy.
# Python
tts = deepgram.TTSv2(model="flux-alexis-en")// TypeScript
const tts = new deepgram.TTSv2({ model: "flux-alexis-en" });Aura 2
Section titled “Aura 2”Deepgram offers 60+ voices across seven languages with Aura 2. Replace the model parameter in TTS with any supported voice:
# Python
tts = deepgram.TTS(model="aura-2-andromeda-en")// TypeScript
const tts = new deepgram.TTS({ model: "aura-2-andromeda-en" });Browse all available voices in the Deepgram voice library.
Go further with Deepgram
Section titled “Go further with Deepgram”- Voices — Deepgram offers 60+ voices across seven languages. Browse the voice library and replace the
modelparameter inTTSwith any supported voice. - Keyterm prompting — Improve recognition of domain-specific vocabulary by passing keyterms to Nova-3 via the STT plugin settings.
- Speaker diarization — Assign a speaker identifier to each word in the transcript using diarization via the STT plugin settings.
- Plugin reference — See the full list of plugin parameters in the LiveKit Deepgram reference for Python and Node.js.