Home
HomeDeepgram developer documentation — APIs, SDKs, and tools for speech-to-text, text-to-speech, voice agents, and audio intelligence.API ReferenceOverview of all Deepgram REST and WebSocket APIs — voice agents, speech-to-text, text-to-speech, audio intelligence, text intelligence, and account management.Voice AgentBuild real-time, interactive voice agents powered by Deepgram’s speech-to-text, LLM integration, and text-to-speech, all over a single WebSocket connection.Speech-to-TextText-to-SpeechIntelligenceAsk AI
Self-Hosted Deployments
Developer Tools
Command-Line Interface
Trust & Security
SDKs
Guides
Fundamentals
Deep Dives
Managing ProjectsUsing Multiple ProjectsWorking With Roles & API ScopesSupported Audio FormatsDeployment OptionsCreating Additional API KeysSafely Removing Team Members With Active API KeysUsing the Sec-WebSocket-ProtocolLogs & Usage DataModel Improvement ProgramWorking With Concurrency Rate LimitsAudio Preprocessing & Barge-InUnderstanding NAT Port Exhaustion
Use Cases
Voice Agent with Pipecat and DeepgramBuild a real-time voice agent with Pipecat for pipeline orchestration, Deepgram Nova-3 for speech-to-text, and Deepgram Aura for text-to-speech.Voice Agent with LiveKit and DeepgramBuild a real-time voice agent with LiveKit for WebRTC transport, Deepgram Nova-3 for speech-to-text, and Deepgram Aura for text-to-speech.Calculate Talk Time Analytics
Integrations
Amazon Connect and DeepgramAWS S3 Presigned URLs and DeepgramAudioCodes (LiveHub) and Deepgram STTGenesys and DeepgramGenesys is a cloud-based platform used by many organizations to manage their call centers. With our plug-and-play Genesys integration, you can have all of your Genesys calls transcribed by Deepgram.LiveKit and DeepgramGoogle Dialogflow CX and DeepgramMake.com and DeepgramPipecat and DeepgramTwilio and Deepgram STTTwilio and Deepgram TTSZapier and DeepgramZoom and Deepgram
Self-Hosted Deployments
General
API Reference
Voice Agent
Agent Configurations
Agent Variables
Speech to Text
Text to Speech
Text Intelligence
Manage
Projects
Models
Keys
Members
Invites
Requests
Usage
Billing
Self-Hosted
Auth
Speak
V2
Audio
Self-Hosted Deployments
Get Started
Build a Voice Agent
Configure
Build
Function Calling
Browser Agent
OverviewFour composable packages that connect your web app to Deepgram's Voice Agent API. Works with any JavaScript framework. Ship in minutes, customize for months.JavaScript SDKAPI reference for @deepgram/agents — the core WebSocket session, microphone capture, and audio playback with volume and frequency analysis for browser-based voice agents.React Hooks & ProviderAPI reference for @deepgram/react — AgentProvider, connection state hooks, playback-aware mode tracking, conversation hooks, component-scoped client tools, and standalone useDeepgramAgent for simpler apps.React UI ComponentsAPI reference for @deepgram/ui — composable components for building voice agent UIs with Deepgram. Includes orb visualizer, waveforms, frequency bars, conversation display, and CSS custom property theming.Widget Embedding GuideHow to embed the Deepgram voice agent widget on any web page. Covers CDN and ES module installation, six layout modes, theming with design tokens, callbacks, and programmatic teardown.
Connect
Telephony Agents
Telephony Integrations
TwilioGenesys Cloud CXLearn how to integrate Genesys Cloud CX with the Deepgram Voice Agent API for real-time conversational AI using the Genesys Audio Connector.Amazon ConnectLearn how to integrate Amazon Connect telephony with Deepgram Voice Agent to build real-time conversational AI with function calling.AudioCodes (LiveHub)
Controls
Inputs: Client Messages
Outputs: Server Events
Optimize
Pre-Recorded Audio
Getting StartedFeature OverviewTemplate AppsGet up and running fast with our pre-recorded speech-to-text template applications, fully integrated with Deepgram out-of-the-box.
Tips and Tricks
Automatically Generating WebVTT & SRT CaptionsAutomatically Transcribe and Summarize Phone CallsGetting Started with Deepgram Whisper CloudGenerating and Saving Transcripts From the TerminalUsing Callbacks to Return Transcripts to Your ServerWhen Callback Is Not ReceivedWhen To Use Multichannel and DiarizationWhen To Use Keywords and Search
Streaming Audio
Conversational STT for Voice Agents (Flux)
Getting StartedFeature OverviewTemplate AppsGet up and running fast with our Flux conversational STT template applications, fully integrated with Deepgram out-of-the-box.End-of-Turn ConfigurationFlux Multilingual & Language PromptingBuild a Flux-enabled Voice AgentWhy Flux's State Machine Matters
Control Messages
Migrating
Tips & Tricks
Transcription (Nova-3)
Getting StartedFeature OverviewLive Streaming Starter KitTemplate AppsGet up and running fast with our streaming speech-to-text template applications, fully integrated with Deepgram out-of-the-box.
Control Messages
Speech Detection
Tips and Tricks
End of Speech Detection While Live StreamingUsing Interim ResultsEndpointing & Interim Results With Live StreamingDetermining Your Audio Format for Live Streaming AudioMeasuring Streaming LatencySTT Troubleshooting WebSocket, NET, and DATA ErrorsRecovering From Connection Errors & Timeouts When Live StreamingUsing Lower-Level Websockets with the Streaming API
Models and Languages
Formatting
Speaker DiarizationDiarize recognizes speaker changes and assigns a speaker to each word in the transcript.DictationFiller WordsFiller Words can help transcribe interruptions in your audio, like "uh" and "um".MeasurementsNumeralsParagraphsParagraphs splits audio into paragraphs to improve transcript readability.Profanity FilteringProfanity Filter looks for recognized profanity and replaces it with asterisks.PunctuationPunctuation adds punctuation and capitalization to your transcript.RedactionRedaction removes sensitive information from your transcripts.Smart FormattingSmart Format can automatically format transcripts to improve readability.Supported Entity TypesUtterancesUtterances segments speech into meaningful semantic units.Utterance Split
Custom Vocabulary
Find and ReplaceFind and Replace searches for terms or phrases in submitted audio and replaces them.Keyterm PromptingKeyterm Prompting allows you to improve Keyword Recall Rate (KRR) for important keyterms or phrases up to 90%.KeywordsKeywords can boost or suppress specialized terminology.SearchSearch searches for terms or phrases in submitted audio.
Media Input Settings
Results Processing
Migrating
Flux TTS
Streaming (WebSocket)
Control Messages
Cross-turn Context & State
Batch (REST)
Aura
Streaming (WebSocket)
Streaming (WebSocket)Getting StartedFeature OverviewTemplate AppsGet up and running fast with our Aura (/v1/speak) streaming text-to-speech template applications, fully integrated with Deepgram out-of-the-box.
Control Messages
Batch (REST)
TTS Voice Controls
Media Output Settings
Results Processing
Tips and Tricks
Real-Time TTS with WebSocketsText Chunking for TTSFormatting text for Aura-2Formatting Text for Aura-2Handling Audio Issues in Text To SpeechSending LLM Outputs to a WebSocketText Chunking for TTS REST OptimizationText to Speech LatencyText to Speech PromptingTTS Troubleshooting WebSocket, NET, and DATA Errors
Self-Hosted Deployments
Audio Intelligence
Feature OverviewEntity DetectionIntent RecognitionIntent Recognition recognizes speaker intent throughout a transcript.Sentiment AnalysisSentiment Analysis recognizes the sentiment throughout an entire transcript.SummarizationSummarization provides a brief summary of the audio.Topic DetectionDetects topics throughout a transcript.
Text Intelligence
Getting StartedFeature OverviewTemplate AppsGet up and running fast with our text intelligence template applications, fully integrated with Deepgram out-of-the-box.Intent RecognitionIntent Recognition recognizes intent throughout the input text.Sentiment AnalysisSentiment Analysis recognizes sentiment throughout the inputed text.SummarizationSummarization provides a brief summary of the input text.Topic DetectionTopic Detection detects topics throughout the input text.
Results Processing
Self-Hosted Deployments
Amazon SageMaker
Get started
Supported ProductsRequesting SageMaker QuotaHow to check and increase AWS service quotas for SageMaker endpoint instance types (ml.g4dn.2xlarge, ml.g5.2xlarge, ml.g6.2xlarge, ml.g6e.2xlarge, ml.g7.2xlarge, ml.g7e.2xlarge) required by Deepgram deployments.Subscribe on AWS MarketplaceSubscribe to a Deepgram SageMaker product through the AWS Marketplace console or API, find the Model Package ARN, and manage private offers.Deploy Deepgram on Amazon SageMakerDeploy with TerraformDeploy a Deepgram SageMaker Endpoint with Terraform using an AWS Marketplace Model Package ARN. Includes IAM role, endpoint configuration, auto-scaling, and optional environment variable overrides.
Configure
Manage endpoints
Invoke a Deepgram SageMaker EndpointValidate a Deepgram SageMaker EndpointTest a Deepgram SageMaker endpoint using the open-source client scripts in the dg-sagemaker repository, with one script per product and language.Update an Amazon SageMaker EndpointRoll out a newer Deepgram Model Package or a different model version on a SageMaker Endpoint that is already serving production traffic.
Auto-Scaling
Monitor & secure
Observability
ObservabilityUse Amazon CloudWatch to monitor Deepgram SageMaker Endpoints. Covers key metrics like ConcurrentRequestsPerModel and FirstChunkLatency, enhanced per-instance metrics, CloudWatch Logs for Deepgram container output, and alarm configuration for proactive alerting.Prometheus & OpenTelemetry MetricsEnable Amazon SageMaker detailed observability on Deepgram Endpoints to collect per-GPU, host, and container Prometheus metrics through an AWS-managed OpenTelemetry Collector, and query them with PromQL from CloudWatch or Prometheus-compatible observability tools.Deepgram Enhanced MetricsReference for the CloudWatch metrics Deepgram SageMaker containers emit via Embedded Metric Format: billing metrics in the Deepgram/SageMakerInference namespace and per-feature usage metrics in the Deepgram/SelfHosted namespace, with dimensions, units, and example queries.
Security and Compliance
Security and ComplianceSecurity and compliance for Deepgram on Amazon SageMaker: TLS requirements for API access, FIPS 140-3 endpoints for the control plane and inference traffic, FedRAMP coverage, network isolation for AWS Marketplace containers, container vulnerability scanning with no Critical or High CVEs, and VPC endpoint options for restricting access to your endpoint.Use FIPS EndpointsHow to configure boto3, the AWS SDKs, and the AWS CLI to reach Deepgram on Amazon SageMaker over FIPS 140-3 endpoints, including bidirectional streaming on port 8443, asynchronous endpoints on s3-fips, the IAM Identity Center caveat, and how to confirm a run used FIPS endpoints.
Docker/Podman
Platform Options
Kubernetes
Platform Options
Deployment
Self Service Licensing & CredentialsDeploy STT ServicesDeploy Flux Model (STT)Flux is a purpose-built, low-latency streaming speech-to-text model tailored for voice agent use cases. This article describes how to ensure Flux is present in your self-hosted Deepgram environment, the configuration steps, and key considerations unique to Flux.Deploy TTS ServicesDeploy Deepgram's TTS services including Aura-2 for conversational AI voice synthesisDeploy Flux TTS ModelDeploy Voice AgentDeploy Deepgram's Voice Agent API for self-hosted conversational AI with real-time speech-to-speech interactionsStatus EndpointCertificate StatusFIPS-Compliant Deployment
Partner Deployment
Modal
Scaling and Deployment Strategies
Scaling and Deployment StrategiesSystem MaintenanceBlue-Green DeploymentAuto-ScalingNetworkingMetrics GuideIngress AuthenticationRedact UsageLog FormatsUsing Private Container RegistriesMirror Deepgram container images to a private registry (AWS ECR, GCP Artifact Registry) for faster autoscaling and tighter network security.