Requesting SageMaker Quota
AWS enforces default service quotas on the number of SageMaker endpoint instances you can run per account per region. Before you can deploy Deepgram on Amazon SageMaker, you may need to request a quota increase for the GPU-accelerated instance types that Deepgram requires.
Instance types used by Deepgram
Section titled “Instance types used by Deepgram”Deepgram SageMaker deployments use the following instance types. Each maps to a separate service quota.
| Instance type | GPU | Quota name |
|---|---|---|
ml.g4dn.2xlarge |
NVIDIA T4 | ml.g4dn.2xlarge for endpoint usage |
ml.g5.2xlarge |
NVIDIA A10G | ml.g5.2xlarge for endpoint usage |
ml.g6.2xlarge |
NVIDIA L4 | ml.g6.2xlarge for endpoint usage |
ml.g6e.2xlarge |
NVIDIA L40S | ml.g6e.2xlarge for endpoint usage |
ml.g7.2xlarge |
NVIDIA RTX PRO 4500 Blackwell | ml.g7.2xlarge for endpoint usage |
ml.g7e.2xlarge |
NVIDIA RTX PRO 6000 Blackwell | ml.g7e.2xlarge for endpoint usage |
Check your current quota
Section titled “Check your current quota”Before requesting an increase, check the quota you already have.
Open Service Quotas
Sign in to the AWS Management Console and navigate to Service Quotas.
Select Amazon SageMaker
In the left-hand menu, select AWS services, then search for and select Amazon SageMaker.
Search for the instance quota
In the search bar, enter the instance type you want to check (for example,
g5.2xlarge). Find the item named similar to ml.g5.2xlarge for endpoint usage.
Run the following command, replacing the --query-text value with the instance type you want to check:
aws service-quotas list-service-quotas \
--service-code sagemaker \
--query "Quotas[?contains(QuotaName, 'ml.g5.2xlarge') && contains(QuotaName, 'endpoint')]" \
--output tableTo check all supported instance types at once:
for INSTANCE in g4dn.2xlarge g5.2xlarge g6.2xlarge g6e.2xlarge g7.2xlarge g7e.2xlarge; do
echo "=== ml.$INSTANCE ==="
aws service-quotas list-service-quotas \
--service-code sagemaker \
--query "Quotas[?contains(QuotaName, 'ml.$INSTANCE') && contains(QuotaName, 'endpoint')].{Name:QuotaName,Value:Value}" \
--output table
doneRequest a quota increase
Section titled “Request a quota increase”If your current quota is 0 or too low for your deployment, submit a quota increase request.
Open the quota detail page
In the Service Quotas console for Amazon SageMaker, search for the instance type (for example,
g5.2xlarge) and select the quota named ml.g5.2xlarge for endpoint usage.Request an increase
Select Request increase at account level.
Enter the new quota value
In the Increase quota value field, enter the number of instances you need. For example, enter
4if you plan to run up to fourml.g5.2xlargeendpoint instances in this region.Submit the request
Select Request. AWS reviews most SageMaker quota requests within a few hours, though some may take several business days.
Repeat for each instance type
If you need quota for additional instance types (
ml.g4dn.2xlarge,ml.g6.2xlarge,ml.g6e.2xlarge,ml.g7.2xlarge,ml.g7e.2xlarge), repeat these steps for each.
Use request-service-quota-increase to submit a request. You need the quota code for each instance type.
First, look up the quota code:
aws service-quotas list-service-quotas \
--service-code sagemaker \
--query "Quotas[?contains(QuotaName, 'ml.g5.2xlarge') && contains(QuotaName, 'endpoint')].{Code:QuotaCode,Name:QuotaName,Value:Value}" \
--output tableThen submit the increase request using the quota code from the previous command:
aws service-quotas request-service-quota-increase \
--service-code sagemaker \
--quota-code <QUOTA_CODE> \
--desired-value 4Replace <QUOTA_CODE> with the code returned by the lookup command, and 4 with the number of instances you need.
To check the status of your request:
aws service-quotas list-requested-service-quota-change-history \
--service-code sagemaker \
--query "RequestedQuotas[?contains(QuotaName, 'g5.2xlarge')].{Status:Status,Desired:DesiredValue,QuotaName:QuotaName}" \
--output tableChoosing the right quota value
Section titled “Choosing the right quota value”The quota value determines how many instances of that type you can run simultaneously. Consider the following when choosing a value:
- Number of Deepgram products — each product (for example, English STT, Spanish STT, TTS) runs as a separate SageMaker endpoint.
- Auto-scaling — if you configure auto-scaling, set the quota high enough to accommodate the maximum instance count across all endpoints.
- Multi-region deployments — quotas are per region. Request increases in every region where you plan to deploy.
Troubleshooting
Section titled “Troubleshooting”| Symptom | Cause | Resolution |
|---|---|---|
ResourceLimitExceeded when creating an endpoint |
The instance quota for the selected type is 0 or fully consumed |
Request a quota increase for that instance type |
Quota request remains PENDING for several days |
AWS is reviewing the request | Open a support case in the AWS Support Center referencing the quota request ID |
| Quota increase approved but endpoint still fails | Quota was increased in a different region | Verify the region in the AWS Management Console matches the region where you are creating the endpoint |