Troubleshooting
If you're experiencing any issues with your Deepgram deployment on Amazon SageMaker AI, start with the Deepgram container logs in Amazon CloudWatch, then work through the common causes below.
View container logs
Section titled “View container logs”If you open the SageMaker AI Endpoint resource details, there will be a link to open the Amazon CloudWatch Log Group for that endpoint. Within the CloudWatch Log Group, there should be a Log Stream that contains the Deepgram logs for all components. You can use the Amazon CloudWatch Logs Live Tail feature to watch logs in near-real-time while you are sending requests to the Deepgram API, via the SageMaker AI APIs.
To use the CloudWatch Logs Live Tail feature locally, from the AWS CLI tool, you can use the following command.
aws logs tail --follow /aws/sagemaker/Endpoints/YOUR_SAGEMAKER_ENDPOINT_NAME --region YOUR_AWS_REGIONEndpoint fails to start (CUDA / driver preflight)
Section titled “Endpoint fails to start (CUDA / driver preflight)”Current Deepgram model packages run a CUDA 13 runtime and require NVIDIA driver 580 or later on the host. If the endpoint boots on an older default inference AMI — which is what happens when the Endpoint Configuration was created from the SageMaker AI console, or without InferenceAmiVersion set — the container fails its CUDA preflight check and the endpoint never reaches InService. The CloudWatch log stream for the endpoint contains a line similar to:
[cuda-preflight] To fix this, set InferenceAmiVersion to al2023-ami-sagemaker-inference-gpu-4-1 on your endpoint configuration's ProductionVariant. That AMI provides NVIDIA driver 580 with CUDA 13.To fix this, create a new Endpoint Configuration with the AWS CLI, Boto3, or Terraform that sets InferenceAmiVersion to al2023-ami-sagemaker-inference-gpu-4-1 on the production variant, then create or update the endpoint with it. See Inference AMI Versions for the available versions and Deepgram's recommendation.
Endpoint is InService but every request returns 400
Section titled “Endpoint is InService but every request returns 400”A 400 on every request usually means the request does not match the product you deployed, not that the endpoint is unhealthy.
- Multilingual Nova-3 listings require
language=multi. Sendinglanguage=en(or any single language code) to a multilingual Nova-3 endpoint returns400, which looks like a dead endpoint. Passlanguage=multiin the query string. - Flux multilingual is selected by model name, not a language parameter. Use
model=flux-general-multi; there is nolanguageparameter for Flux multilingual. - Streaming-mode bundles reject synchronous invocation. A product listing published for streaming returns
400 No such model/language/tierwhen called through the synchronous/invocationspath (InvokeEndpoint). UseInvokeEndpointWithBidirectionalStreamfor streaming listings, or deploy the batch listing for synchronous invocation. See Invoke a Deepgram SageMaker Endpoint. - Streaming clients see a 424, not the 400. When the container rejects a bidirectional streaming request, the streaming client receives HTTP
424withFailed to establish WebSocket connection(for example,ModelStreamError). The underlying400and its reason are only visible in the endpoint's CloudWatch logs (see View container logs). The causes are the same as above: a multilingual listing called withoutlanguage=multi, Flux multilingual called with alanguageparameter instead ofmodel=flux-general-multi, or a model, language, or tier the deployed listing does not include.
Endpoint stuck in Creating
Section titled “Endpoint stuck in Creating”If the endpoint stays in Creating well beyond the usual several minutes, or moves to Failed, check FailureReason:
aws sagemaker describe-endpoint --endpoint-name YOUR_SAGEMAKER_ENDPOINT_NAME --query '{Status:EndpointStatus,FailureReason:FailureReason}'Common causes:
ModelDataDownloadTimeoutInSecondsis too low. Large multilingual bundles can take longer than the default download window. Recreate the Endpoint Configuration with a higher value on the production variant (Deepgram's CLI steps use600; large multilingual Nova-3 bundles may need1800).- Capacity or quota. The instance type is unavailable in the Availability Zone, or your account quota for it is
0. AResourceLimitExceededfailure reason points to quota — see Requesting SageMaker Quota. With an instance pool, every type in the pool needs quota; quota does not fall back between types. - Provisioning failure with no container logs. The endpoint goes
Failedafter roughly 3 minutes withFailureReasonRequest to service failed. If failure persists after retry, contact customer support., and no/aws/sagemaker/Endpoints/<endpoint-name>log group is ever created. The container never started, so this is a provisioning-level problem (typically GPU capacity for that instance type in that zone), not a problem with the model image. Retry with an ordered instance pool so SageMaker can fall back to another type, or deploy in another region.
Checklist
Section titled “Checklist”If you experience any issues using Deepgram services running on the Amazon SageMaker AI platform, please review this checklist before contacting Deepgram support.
- Ensure that your application's AWS IAM User or IAM Role has permission to call the
InvokeEndpointWithBidirectionalStreamSageMaker AI action. - Ensure your application is targeting the correct AWS account and region, where your SageMaker Endpoint exists.
- Ensure the Deepgram product you've deployed (eg. streaming Speech-to-Text), from the AWS Marketplace, corresponds to the Deepgram API you're calling.
- Ensure the Endpoint Configuration pins
InferenceAmiVersiontoal2023-ami-sagemaker-inference-gpu-4-1. The SageMaker AI console does not expose this setting; create the Endpoint Configuration with the AWS CLI or API, or with Terraform. See Inference AMI Versions. - If you subscribed through a private offer in an AWS organization, ensure the linked account deploying the endpoint has accepted the offer or holds a License Manager entitlement. See Private offers.