Deploy Speech-to-Text (STT) Services
This guide covers deploying Deepgram’s Speech-to-Text (STT) services for transcription and real-time speech recognition capabilities.
Deepgram STT services use the same container images (quay.io/deepgram/self-hosted-api and quay.io/deepgram/self-hosted-engine) as TTS deployments, as described in the Deployment Environments overview. However, Deepgram strongly recommends configuring each node for a specific use case—either STT or TTS—for optimal performance and resource utilization.
We need to download and deploy these images from a container image repository, along with STT-specific configuration files and AI models that will be provided by Deepgram.
Prerequisites
Section titled “Prerequisites”Before you begin, you will need to complete the Deployment Environments guide, as well as all sub-guides to complete your environment configuration.
You will also need to complete the Self Service Licensing & Credentials guide to authenticate your products with Deepgram’s licensing servers and pull Deepgram container images from Quay.
Get Deepgram Products
Section titled “Get Deepgram Products”Cache Container Image Repository Credentials
Section titled “Cache Container Image Repository Credentials”Use the image repository credentials you generated in the self-service licensing and credentials guide to login to Quay on your deployment environment. Once your credentials are cached locally, you should not have to log in again (until after you manually log out).
Choosing A Deployment Type
Section titled “Choosing A Deployment Type”For customers deploying Deepgram’s self-hosted solution in highly available production environments, Deepgram recommends the License Proxy, which is a caching proxy that communicates with the Deepgram-hosted license server to ensure uptime and simplify network security. If you aren’t certain which products your contract includes or if you are interested in adding the License Proxy to your self-hosted deployment, please consult your Deepgram Account Representative.
If your project is authorized to use the License Proxy, in the following section you may choose to use either a standard deployment or a deployment with the License Proxy.
See the License Proxy guide for more details on the benefits and setup.
Import Your Docker Compose, Container Configuration, and Model Files
Section titled “Import Your Docker Compose, Container Configuration, and Model Files”Before you can run your self-hosted deployment, you must configure the required components. To do this, you will need to customize your configuration files and create a directory to house models that have been encrypted for use in your requests.
-
In your home directory, create the following directories. This is where you will save your Docker Compose files, Deepgram configuration files, and Deepgram model files.
Bash -
The following sub-steps will download Deepgram configuration template files:
-
Choose to use either a standard deployment, or a deployment with the License Proxy if it is enabled for your project. Execute one of the following lines accordingly:
Bash -
Download the appropriate files from the self-hosted resources repository.
Bash -
Modify the Docker/Podman templates to point to the correct paths.
Bash
-
-
Download the Deepgram model files (file extension
.dg) that have been provided to you by your Deepgram Account Representative.-
Create a fresh text file and copy over the list of links to the provided models.
Bash After editing, your
model_links.txtfile should look like this:https://LINK_TO_MODEL_1.dg https://LINK_TO_MODEL_2.dg https://LINK_TO_MODEL_3.dg ... https://LINK_TO_MODEL_N.dg -
Download all the models specified in your
model_links.txtfile.Bash -
Modify the Compose file templates to point to your models directory.
Bash
-
See the Model Maintenance guide for more details on how Deepgram models power voice AI inference in your self-hosted deployment.
Customize Your Configuration
Section titled “Customize Your Configuration”Once you have downloaded all provided files to your deployment machine, you need to update your configuration for your specific deployment environment.
Credentials
Section titled “Credentials”You will need to have an environment variable DEEPGRAM_API_KEY exported with your self-hosted API key secret. See our Self Service Licensing & Credentials guide for instructions on generating a self-hosted API key for use in this section.
Configuration Files
Section titled “Configuration Files”Compose File
Section titled “Compose File”The Docker Compose or Podman Compose configuration file makes it possible to spin up the containers using a single command. This makes spinning up a standard POC deployment quick and easy.
Make sure to export your self-hosted API key secret in your deployment environment.
api.toml , engine.toml, and license-proxy.toml
Section titled “api.toml , engine.toml, and license-proxy.toml”The API and Engine containers, and the optional License Proxy container, are configured with TOML configuration files. The templates provided by Deepgram in the deepgram-self-hosted repository contain sane defaults that will work well for most use cases; these need to be mounted to the containers (see the volumes section in the Compose file).
There are header comments describing each config value available in both of these files. If you have any questions about modifying these files, refer to those comments or reach out to your Deepgram Account Representative.
Testing Your Containers
Section titled “Testing Your Containers”To make sure your Deepgram self-hosted deployment is properly configured and running, you will want to run the containers and make a sample request.
Start the Deepgram Containers
Section titled “Start the Deepgram Containers”Now that you have your configuration files and AI models set up and in the correct location to be used by the container, use Docker Compose to run the container:
You can then view the running containers with the container process status command, and optionally view the logs of each container to verify their status.
Networking Considerations
Section titled “Networking Considerations”If you are running your API and Engine nodes on separate instances, you may need to add an inbound rule for port 8080 to the API instances’ security group, so that port 8080 is reachable from where you are initiating your requests.
Unless you have HTTPS or TLS running on your API instance, construct your Deepgram API endpoint with http://, not https://, and ws://, not wss:// (for instance, http://localhost:8080/v1/listen).
Test Your Deepgram Setup with a Sample Request
Section titled “Test Your Deepgram Setup with a Sample Request”Test your environment and container setup with a local file.
-
Download a sample file from Deepgram (or supply your own file).
Bash -
Send your audio file to your local Deepgram setup for transcription.
Bash
You should receive a JSON response with the transcription and associated metadata. Congratulations - your self-hosted setup is working!
What’s Next
Now that you have a basic Deepgram setup working, take some time to learn about building up to a production-level environment, as well as helpful Deepgram add-on services.