A reference implementation for building a secure inbound telephony voice agent using Deepgram's Voice Agent API and Twilio. Uses Deepgram Flux for speech-to-text with native turn-taking optimized for real-time voice agent conversations. Includes webhook endpoint protection and Twilio request signature validation out of the box.
Callers dial a phone number and talk to an AI receptionist that can check appointment availability, book appointments, look up existing appointments, and cancel appointments, all through natural voice conversation. While this example is specific to a dental office, it's a great jumping off point for any inbound use case – simply edit the prompts and function calls accordingly.
Single WebSocket bridge: The core of the system is VoiceAgentSession, which bridges two WebSocket connections (one to Twilio, one to Deepgram). It translates between Twilio's JSON-based protocol and Deepgram's binary audio protocol.
Barge-in: When the Deepgram Voice Agent detects that the user started speaking, the server sends a Twilio "clear" event to immediately stop playing agent audio. This prevents the agent from talking over the caller.
Function calls: The Deepgram Voice Agent API supports tool use. When the LLM decides to call a function (like checking appointment availability), Deepgram sends a function call event. The server executes it against the backend service and sends the result back to Deepgram, which incorporates it into the agent's next response.
Protocol-agnostic server: The server doesn't know whether audio comes from a real Twilio call or from the local development client (dev_client.py). Both send identical WebSocket messages. This means you can develop and test without a phone or Twilio account.
For a deeper look at the call flow, session lifecycle, and component details, see docs/ARCHITECTURE.md.
Depending on your region, Twilio may require address verification before you can purchase a phone number. The setup wizard will surface any errors from Twilio. Follow the instructions in the Twilio console if prompted.
If you prefer to run the server locally instead of deploying to Fly.io, you can use a tunnel to expose your locally running voice agent server via a public URL.
You can test the full voice agent conversation without a phone or Twilio account using dev_client.py, which connects to the server over WebSocket and streams audio from your microphone.
Agent: "Thank you for calling Bright Smile Dental! How can I help you today?"
Caller: "Hi, I'd like to schedule a cleaning."
Agent: "I'd be happy to help you schedule a cleaning. Do you have a preferred
date or day of the week?"
Caller: "How about next Tuesday?"
[Agent calls check_available_slots(date="2026-03-03")]
Agent: "I have a few openings next Tuesday with Lisa Thompson, our hygienist.
There's a slot at 10 AM and another at 2 PM. Which would work better
for you?"
Caller: "10 AM works."
Agent: "Great, a cleaning with Lisa Thompson on Tuesday March 3rd at 10 AM.
Can I get your name and phone number to book that?"
Caller: "Sure, it's Maria Garcia, 555-0199."
Agent: "I have a cleaning with Lisa Thompson on Tuesday March 3rd at 10 AM
for Maria Garcia. Shall I go ahead and book that?"
Caller: "Yes please."
[Agent calls book_appointment(patient_name="Maria Garcia",
patient_phone="555-0199",
slot_id="slot-abc123")]
Agent: "You're all set! Your cleaning is booked for Tuesday March 3rd at
10 AM with Lisa Thompson. Is there anything else I can help with?"
Caller: "No, that's all. Thanks!"
Agent: "You're welcome! We'll see you on Tuesday. Have a great day!"
[Agent calls end_call(reason="appointment_booked")]
Edit the SYSTEM_PROMPT in voice_agent/agent_config.py. This is where you define who the agent is, what it knows, and how it behaves. See docs/PROMPT_GUIDE.md for voice-specific prompt best practices.
This reference implementation is configured for English using Deepgram Flux (flux-general-en) for STT, which provides native turn-taking optimized for voice agents. To add multilingual support, switch the STT model to flux-general-multi and set language_hints to an array of BCP-47 codes to bias toward expected languages. For building voice agents in other languages, see the Deepgram multilingual voice agent guide.