Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Deploy Deepgram on Amazon SageMaker

This guide deploys a Deepgram AWS Marketplace Model Package as a SageMaker AI Endpoint using the AWS CLI or the AWS SDK for Python (Boto3). The SageMaker Endpoint resource represents the compute instances that run the Deepgram Voice AI services. For an overview of running Deepgram on SageMaker, including benefits, tradeoffs, and pricing, see Amazon SageMaker.

Deploy a real-time endpoint. It serves both live streaming (InvokeEndpointWithBidirectionalStream) and synchronous single-file transcription (InvokeEndpoint, up to 25 MB per request).

Deploy on an ordered instance pool rather than a single instance type. A single type has no fallback: when AWS is short of that GPU in the Availability Zone, the endpoint goes Failed with Request to service failed a few minutes in, or InsufficientInstanceCapacity, and this happens routinely for popular GPU types. With instance pools, SageMaker tries each type in priority order and falls back to the next when one is capacity-constrained.

Order the pool as follows:

  1. The listing’s recommended type first (for example ml.g6.2xlarge for Speech-to-Text). It is the type Deepgram validated the model on and the best price/performance.
  2. Same-or-newer generation with similar per-instance capacity next (g6 → g6e → g7). Keeping capacity similar matters if you auto-scale, because the predefined scaling metrics are per instance and do not account for a mixed fleet.
  3. Older generations last, as insurance (g5, and g4dn where supported).
  4. Never include a type the product does not support: g4dn for Flux, g5/g4dn for Flux TTS, or any single-GPU type for Aura-2. See Instance types.
  5. Up to 5 types. Three is the sweet spot.

VariantInstanceProvisionTimeoutInSeconds is the per-type wait: SageMaker tries each type for that long before moving to the next. 300 is recommended (AWS allows 60–3600), so a three-type pool can stay in Creating for up to about 15 minutes before it fails.

Prefer a single instance type only for a stated reason: a Machine Learning Savings Plan or reservation on that type, or an auto-scaling concurrency target you measured on a specific GPU.

SageMaker assumes an execution role to run the Model Package on your behalf. You only need to create a single SageMaker execution role, and can reuse this IAM Role to deploy multiple SageMaker Endpoints.

Bash
Python
import json
import boto3

iam = boto3.client("iam")

role = iam.create_role(
    RoleName="deepgram-sagemaker-execution",
    AssumeRolePolicyDocument=json.dumps({
        "Version": "2012-10-17",
        "Statement": [{
            "Effect": "Allow",
            "Principal": {"Service": "sagemaker.amazonaws.com"},
            "Action": "sts:AssumeRole",
        }],
    }),
)
iam.attach_role_policy(
    RoleName="deepgram-sagemaker-execution",
    PolicyArn="arn:aws:iam::aws:policy/AmazonSageMakerFullAccess",
)
execution_role_arn = role["Role"]["Arn"]

Note

A newly created IAM role can take around 10 seconds to become assumable. If CreateModel fails with Could not assume role immediately after create-role, the error is transient — wait a few seconds and retry.

  1. Set variables

    Choose names for the three SageMaker resources, and set the Model Package ARN and execution role ARN.

    • MODEL_PACKAGE_ARN identifies the Deepgram product version and AWS Region you subscribed to. It is region-specific, so copy the ARN for the Region you deploy in. To find it, open the AWS Marketplace Manage subscriptions console, click Configure on your Deepgram subscription, choose AWS command line interface (CLI) under Service, select the product version, and copy the ARN for your Region from the Model ARNs list. See Find the Model Package ARN for the full steps.
    • EXECUTION_ROLE_ARN is the role you created in Create an IAM execution role.
    • The instance pool is set in the Endpoint Configuration step. ml.g6.2xlarge is the recommended first type for Speech-to-Text; see Choose instance types and Instance types for Text-to-Speech and the other supported families.
    AWS CLI
    Boto3
    import boto3
    
    AWS_REGION = "us-east-1"
    MODEL_NAME = "deepgram-streaming-stt"
    ENDPOINT_CONFIG_NAME = "deepgram-streaming-stt-config"
    ENDPOINT_NAME = "my-deepgram-streaming-stt"
    MODEL_PACKAGE_ARN = "arn:aws:sagemaker:us-east-1:123456789012:model-package/deepgram-stt-nova-3/1"
    EXECUTION_ROLE_ARN = "arn:aws:iam::123456789012:role/deepgram-sagemaker-execution"
    
    sagemaker = boto3.client("sagemaker", region_name=AWS_REGION)
  2. Create the Model

    The SageMaker Model wraps the Marketplace Model Package and the execution role.

    Bash
    Python
    sagemaker.create_model(
        ModelName=MODEL_NAME,
        ExecutionRoleArn=EXECUTION_ROLE_ARN,
        PrimaryContainer={"ModelPackageName": MODEL_PACKAGE_ARN},
        EnableNetworkIsolation=True,
    )

    EnableNetworkIsolation=true is mandatory for AWS Marketplace model packages — SageMaker rejects the Model otherwise. Network isolation is also why the container cannot reach external services; see Limitations.

    To pass DEEPGRAM_API_* or DEEPGRAM_ENGINE_* configuration overrides, add an Environment map to the container definition. See Configure Amazon SageMaker Deployments.

    EnableNetworkIsolation=true is mandatory for AWS Marketplace model packages — SageMaker rejects the Model otherwise. Network isolation is also why the container cannot reach external services; see Limitations.

    To pass DEEPGRAM_API_* or DEEPGRAM_ENGINE_* configuration overrides, add an Environment map to the container definition. See Configure Amazon SageMaker Deployments.

  3. Create the Endpoint Configuration

    The Endpoint Configuration sets the instance pool, instance count, and — critically — the host inference AMI version the instances boot with.

    The examples use an ordered instance pool, as recommended in Choose instance types. Adjust the types and order for your product.

    Instance pool (recommended)

    To deploy on a single instance type instead (see when to prefer a single type), replace InstancePools and VariantInstanceProvisionTimeoutInSeconds with InstanceType:

    Single instance type

    Instance pool (recommended)

    Instance pool (recommended)
    sagemaker.create_endpoint_config(
        EndpointConfigName=ENDPOINT_CONFIG_NAME,
        ProductionVariants=[{
            "VariantName": "AllTraffic",
            "ModelName": MODEL_NAME,
            "InitialInstanceCount": 1,
            "InstancePools": [
                {"InstanceType": "ml.g6.2xlarge", "Priority": 1},
                {"InstanceType": "ml.g6e.2xlarge", "Priority": 2},
                {"InstanceType": "ml.g5.2xlarge", "Priority": 3},
            ],
            "VariantInstanceProvisionTimeoutInSeconds": 300,
            "InferenceAmiVersion": "al2023-ami-sagemaker-inference-gpu-4-1",
            "ModelDataDownloadTimeoutInSeconds": 600,
            "ContainerStartupHealthCheckTimeoutInSeconds": 300,
        }],
    )

    To deploy on a single instance type instead (see when to prefer a single type), replace InstancePools and VariantInstanceProvisionTimeoutInSeconds with InstanceType:

    Single instance type

    Single instance type
    sagemaker.create_endpoint_config(
        EndpointConfigName=ENDPOINT_CONFIG_NAME,
        ProductionVariants=[{
            "VariantName": "AllTraffic",
            "ModelName": MODEL_NAME,
            "InitialInstanceCount": 1,
            "InstanceType": "ml.g6.2xlarge",
            "InferenceAmiVersion": "al2023-ami-sagemaker-inference-gpu-4-1",
            "ModelDataDownloadTimeoutInSeconds": 600,
            "ContainerStartupHealthCheckTimeoutInSeconds": 300,
        }],
    )

    Keep VariantName=AllTraffic: the Update an Amazon SageMaker Endpoint procedure and the Terraform configuration use the same variant name.

    ModelDataDownloadTimeoutInSeconds and ContainerStartupHealthCheckTimeoutInSeconds are ceilings, not fixed waits: they set how long SageMaker allows the model package to download and the container to load models before it marks the endpoint failed. 600 and 300 suit most products. Large multilingual Nova-3 bundles may need ModelDataDownloadTimeoutInSeconds of 1800.

    Keep VariantName=AllTraffic: the Update an Amazon SageMaker Endpoint procedure and the Terraform configuration use the same variant name.

    ModelDataDownloadTimeoutInSeconds and ContainerStartupHealthCheckTimeoutInSeconds are ceilings, not fixed waits: they set how long SageMaker allows the model package to download and the container to load models before it marks the endpoint failed. 600 and 300 suit most products. Large multilingual Nova-3 bundles may need ModelDataDownloadTimeoutInSeconds of 1800.

  4. Create the Endpoint

    AWS CLI
    Boto3
    sagemaker.create_endpoint(
        EndpointName=ENDPOINT_NAME,
        EndpointConfigName=ENDPOINT_CONFIG_NAME,
    )
  5. Wait for InService

    It takes several minutes for the endpoint to download the model package, start the container, and pass its health check.

    Bash
    Python
    sagemaker.get_waiter("endpoint_in_service").wait(EndpointName=ENDPOINT_NAME)
    
    status = sagemaker.describe_endpoint(EndpointName=ENDPOINT_NAME)["EndpointStatus"]
    print(status)  # InService

    If the endpoint moves to Failed or stays in Creating, see Troubleshooting.

    If the endpoint moves to Failed or stays in Creating, see Troubleshooting.

  6. Verify

    Send a first request to confirm the endpoint transcribes audio — see Validate a Deepgram SageMaker Endpoint. For the streaming and synchronous invocation APIs and the Deepgram SDK SageMaker transport, see Invoke a Deepgram SageMaker Endpoint.

A SageMaker Endpoint Configuration can pin an inference AMI version — the SageMaker-managed host image supplying the NVIDIA driver and container runtime your instances boot with. It is independent of the Deepgram container: it determines which driver the container runs against. If you do not set it, SageMaker selects a default for your instance type, which on older GPU families is an older driver.

AMI version NVIDIA driver CUDA
al2-ami-sagemaker-inference-gpu-2 535 12.2
al2-ami-sagemaker-inference-gpu-2-1 535 12.2
al2-ami-sagemaker-inference-gpu-3-1 550 12.4
al2023-ami-sagemaker-inference-gpu-4-1 580 13.0

For the full list of AMI versions and their driver and CUDA versions, see InferenceAmiVersion in the SageMaker API reference. For the driver each instance family runs by default, see the SageMaker GPU driver table.

The CLI and Boto3 steps above already pin InferenceAmiVersion to al2023-ami-sagemaker-inference-gpu-4-1 on the production variant. Terraform users set the same value through the inference_ami_version variable — see Deploy with Terraform. The SageMaker AI console does not expose this setting, which is why the console path below is not recommended.

Section titled “Deploy with the SageMaker AI console (not recommended)”
Console steps
  1. In the AWS Management Console, navigate to the AWS Marketplace Manage subscriptions console

  2. On the Active subscriptions tab, find the subscription for the Deepgram product you want to deploy (eg. Deepgram Voice AI- Nova-3 Monolingual Speech-to-Text (STT) Streaming)

  3. Click the Configure button in the Actions column on the right-hand side

  4. In the Setup box, under Service, choose Amazon SageMaker AI console

  5. Under the Version header, select the product version from the dropdown. If the listing has more than one version, read the version name and the release notes to understand the set of languages (or features) each version provides, and choose the version that matches your needs

  6. Select the AWS Region you want to deploy to

  7. Under Amazon SageMaker options, keep Create real-time inference endpoint selected

  8. Click the Create endpoint button. You’ll be redirected to the Amazon SageMaker AI console

  9. Provide a name for the model (eg. deepgram-streaming-stt)

  10. Under IAM Role, select the SageMaker execution role that you created

  11. Click the Next button

  12. Provide an Endpoint Name, such as my-deepgram-streaming-stt

  13. Leave Async invocation config turned off. Asynchronous endpoints are temporarily not supported for Marketplace-hosted Deepgram.

  14. Under Variants ➡️ Production, scroll all the way to the right, and click Edit

  15. If desired, select Choose Other Instance Type and select the instance type you want to deploy to (eg. ml.g6.2xlarge), then click Save

  16. Click the Create Endpoint Configuration button

  17. Click the Submit button, to create the SageMaker AI Endpoint

After following these steps, you should see a new Endpoint in your AWS account. If you don’t see the Endpoint, ensure that you have selected the correct AWS region in the AWS Management Console. It may take several minutes for the Endpoint to change to status InService. Once the Endpoint status has changed to InService, you can monitor the Amazon CloudWatch Logs for the Endpoint to ensure normal operation of the Deepgram services.

Delete the three resources in reverse order. Billing for SageMaker compute and Deepgram usage stops when the endpoint is deleted; your AWS Marketplace subscription remains active and can be reused for the next deployment.

AWS CLI
Boto3
sagemaker.delete_endpoint(EndpointName=ENDPOINT_NAME)
sagemaker.delete_endpoint_config(EndpointConfigName=ENDPOINT_CONFIG_NAME)
sagemaker.delete_model(ModelName=MODEL_NAME)
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu