Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Health Checks & Automatic Recovery

Amazon SageMaker polls a /ping endpoint on every instance backing your Endpoint. That response decides whether the instance receives inference requests, and whether SageMaker replaces it.

Deepgram containers report healthy only when they can actually serve inference — models loaded, inference path functional — rather than answering a static 200. An instance that has failed to load its model reports that state instead of silently collecting requests it cannot serve.

A Deepgram Endpoint downloads its model artifacts and loads them into GPU memory before serving anything. For larger bundles this takes several minutes. Throughout, /ping returns 503, the Endpoint stays in Creating, and SageMaker routes no traffic. The container logs the reason at INFO:

/ping returning 503 — system still initializing (expected during container startup; not yet ready to serve)

Once the models load, /ping returns 200, the Endpoint moves to InService, and SageMaker begins routing.

SageMaker allows a bounded window for a new instance to start passing health checks. Miss it and the instance launch fails, leaving the Endpoint in a failed state. Two fields on the production variant extend that window for a large bundle or a slow artifact download:

Field Purpose
ModelDataDownloadTimeoutInSeconds Time allowed to download model artifacts from Amazon S3.
ContainerStartupHealthCheckTimeoutInSeconds Time allowed for the container to begin passing /ping health checks.

Both accept up to 3600 seconds. They are ceilings, not delays — raising them does not slow a healthy deployment.

SageMaker keeps polling /ping every few seconds. If the container reaches a state where it should not serve — the inference engine stops responding, for example — it reports unhealthy, so SageMaker stops routing new traffic to it and, if the fault persists, replaces it.

Reporting unhealthy is not the same as refusing requests. A container that is simply busy — at its configured stream limit, for example — reports unhealthy while still passing requests through to the inference API: one that fits is served normally, and one that does not receives the API’s own response rather than a generic container error. The container refuses requests itself only when something is genuinely wrong and shedding load is the fastest way back to health.

SageMaker then replaces the instance automatically. Replacement is not instantaneous: AWS requires a sustained failure signal, so a brief blip does not cycle your fleet.

/ping is a yes/no signal consumed by SageMaker. The container also publishes the same health state as a Prometheus gauge, sagemaker_endpoint_health, so you can chart it and alarm on it. With detailed observability enabled, it reaches CloudWatch automatically.

All four series are always present. The current state reports 1, the rest report 0:

sagemaker_endpoint_health{state="healthy"} 1
sagemaker_endpoint_health{state="initializing"} 0
sagemaker_endpoint_health{state="degraded"} 0
sagemaker_endpoint_health{state="critical"} 0
state Meaning What to do
initializing Normal startup while models load. /ping returns 503 and SageMaker routes no traffic. Nothing. Expect this on every new instance.
healthy The container is serving inference. Nothing.
degraded A recoverable fault, or a container at its stream limit. /ping reports unhealthy; requests that the container can serve are still served. Watch it. A brief degraded that returns to healthy is self-recovery working as designed. Sustained degraded under heavy load usually means the endpoint needs more capacity, not repair.
critical The container cannot recover on its own. Only replacement clears this state. SageMaker replaces the instance. If it recurs, contact your Deepgram representative with the Endpoint name and Region.

Because every state is always emitted, an alarm on any one of them never reads “no data” while the container is running. A series that disappears entirely means the scrape failed — a different condition, worth alarming on separately.

critical is deliberately slow to arrive. A container only reaches it after the inference engine has looked unhealthy continuously for several minutes, so that a container which is simply at its stream limit sheds load and recovers instead of being written off. Any sign of recovery in that window resets the clock.

That means critical is a reliable signal but not an early one. For early warning, use the companion gauge:

sagemaker_time_to_critical_seconds 214

It counts down the seconds remaining before the container would enter critical. It reports -1 whenever nothing is counting — the normal reading on a healthy container — and 0 once the container has entered critical.

Guard alarms against that -1, or they fire constantly on healthy containers:

sagemaker_time_to_critical_seconds >= 0 and sagemaker_time_to_critical_seconds < 120

A container that dips into a countdown and returns to -1 recovered on its own, exactly as intended.

/ping reports instance health. Each bidirectional streaming connection has its own liveness check defined by the WebSocket protocol (RFC 6455): SageMaker sends a Ping frame about once a minute, the container replies with a Pong, and several consecutive unanswered Pings close that connection.

The two are independent — a closed connection does not mean the instance is unhealthy, so client applications should reconnect on an unexpected close. See Container Contract to Support Bidirectional Streaming Capabilities.

These appear in the Endpoint’s CloudWatch Log Group, /aws/sagemaker/Endpoints/YOUR_ENDPOINT_NAME. See Observability for Amazon SageMaker for how to read and filter them.

Log pattern Meaning
/ping returning 503 — system still initializing Logged at INFO. Normal during startup while models load. It stops once the container is ready.
INFO Deepgram Engine is ready The inference engine has loaded models and is accepting requests.
composite health check failed Not expected in normal operation. Contact your Deepgram representative to help troubleshoot. Include your Endpoint name, AWS Region, and the surrounding log lines.
What you observe What to do
Endpoint stays in Creating, then Failed Check the container logs — the reason is in the container output, not the SageMaker console error. If the container was still loading when the window expired, raise the two timeout fields above.
An instance was replaced once and service recovered Nothing. Automatic recovery worked.
Instances are replaced repeatedly, or composite health check failed appears Contact your Deepgram representative with the Endpoint name, Region, and log excerpts.
sagemaker_endpoint_health{state="critical"} is 1 The instance will not self-recover. Let SageMaker replace it. Contact your Deepgram representative if it recurs — a restart alone will not fix the underlying fault.
sagemaker_endpoint_health{state="degraded"} is 1 briefly, then healthy Nothing. Self-recovery worked.

To catch this before your users do, alarm on Invocation5XXErrors — see Configure CloudWatch alarms.


Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu