Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Requesting SageMaker Quota

AWS enforces default service quotas on the number of SageMaker endpoint instances you can run per account per region. Before you can deploy Deepgram on Amazon SageMaker, you may need to request a quota increase for the GPU-accelerated instance types that Deepgram requires.

Deepgram SageMaker deployments use the following instance types. Each maps to a separate service quota.

Instance type GPU Quota name
ml.g4dn.2xlarge NVIDIA T4 ml.g4dn.2xlarge for endpoint usage
ml.g5.2xlarge NVIDIA A10G ml.g5.2xlarge for endpoint usage
ml.g6.2xlarge NVIDIA L4 ml.g6.2xlarge for endpoint usage
ml.g6e.2xlarge NVIDIA L40S ml.g6e.2xlarge for endpoint usage
ml.g7.2xlarge NVIDIA RTX PRO 4500 Blackwell ml.g7.2xlarge for endpoint usage
ml.g7e.2xlarge NVIDIA RTX PRO 6000 Blackwell ml.g7e.2xlarge for endpoint usage

Before requesting an increase, check the quota you already have.

  1. Open Service Quotas

    Sign in to the AWS Management Console and navigate to Service Quotas.

  2. Select Amazon SageMaker

    In the left-hand menu, select AWS services, then search for and select Amazon SageMaker.

  3. Search for the instance quota

    In the search bar, enter the instance type you want to check (for example, g5.2xlarge). Find the item named similar to ml.g5.2xlarge for endpoint usage.

Run the following command, replacing the --query-text value with the instance type you want to check:

Bash
aws service-quotas list-service-quotas \
  --service-code sagemaker \
  --query "Quotas[?contains(QuotaName, 'ml.g5.2xlarge') && contains(QuotaName, 'endpoint')]" \
  --output table

To check all supported instance types at once:

Bash
for INSTANCE in g4dn.2xlarge g5.2xlarge g6.2xlarge g6e.2xlarge g7.2xlarge g7e.2xlarge; do
  echo "=== ml.$INSTANCE ==="
  aws service-quotas list-service-quotas \
    --service-code sagemaker \
    --query "Quotas[?contains(QuotaName, 'ml.$INSTANCE') && contains(QuotaName, 'endpoint')].{Name:QuotaName,Value:Value}" \
    --output table
done

If your current quota is 0 or too low for your deployment, submit a quota increase request.

  1. Open the quota detail page

    In the Service Quotas console for Amazon SageMaker, search for the instance type (for example, g5.2xlarge) and select the quota named ml.g5.2xlarge for endpoint usage.

  2. Request an increase

    Select Request increase at account level.

  3. Enter the new quota value

    In the Increase quota value field, enter the number of instances you need. For example, enter 4 if you plan to run up to four ml.g5.2xlarge endpoint instances in this region.

  4. Submit the request

    Select Request. AWS reviews most SageMaker quota requests within a few hours, though some may take several business days.

  5. Repeat for each instance type

    If you need quota for additional instance types (ml.g4dn.2xlarge, ml.g6.2xlarge, ml.g6e.2xlarge, ml.g7.2xlarge, ml.g7e.2xlarge), repeat these steps for each.

Use request-service-quota-increase to submit a request. You need the quota code for each instance type.

First, look up the quota code:

Bash
aws service-quotas list-service-quotas \
  --service-code sagemaker \
  --query "Quotas[?contains(QuotaName, 'ml.g5.2xlarge') && contains(QuotaName, 'endpoint')].{Code:QuotaCode,Name:QuotaName,Value:Value}" \
  --output table

Then submit the increase request using the quota code from the previous command:

Bash
aws service-quotas request-service-quota-increase \
  --service-code sagemaker \
  --quota-code <QUOTA_CODE> \
  --desired-value 4

Replace <QUOTA_CODE> with the code returned by the lookup command, and 4 with the number of instances you need.

To check the status of your request:

Bash
aws service-quotas list-requested-service-quota-change-history \
  --service-code sagemaker \
  --query "RequestedQuotas[?contains(QuotaName, 'g5.2xlarge')].{Status:Status,Desired:DesiredValue,QuotaName:QuotaName}" \
  --output table

The quota value determines how many instances of that type you can run simultaneously. Consider the following when choosing a value:

  • Number of Deepgram products — each product (for example, English STT, Spanish STT, TTS) runs as a separate SageMaker endpoint.
  • Auto-scaling — if you configure auto-scaling, set the quota high enough to accommodate the maximum instance count across all endpoints.
  • Multi-region deployments — quotas are per region. Request increases in every region where you plan to deploy.
Symptom Cause Resolution
ResourceLimitExceeded when creating an endpoint The instance quota for the selected type is 0 or fully consumed Request a quota increase for that instance type
Quota request remains PENDING for several days AWS is reviewing the request Open a support case in the AWS Support Center referencing the quota request ID
Quota increase approved but endpoint still fails Quota was increased in a different region Verify the region in the AWS Management Console matches the region where you are creating the endpoint
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu