Configure Modal Resources
With Modal, hardware resources and autoscaling configuration are specified alongside your application code. Update the paraameters in this section by editing the values in app.py and redeploying.
When you clone the repo, the values are configured for an STT deployment in us-west.
# modal_deepgram/app.py
@app.cls(
image=engine_base_image.env({"DEPLOY_LABEL": DEPLOY_LABEL}),
volumes={
MODELS_PATH: models_vol,
CACHE_PATH: cache_vol,
},
gpu="L4",
secrets=[modal.Secret.from_name("deepgram")],
timeout=30 * MINUTES,
cpu=4,
memory=32 * 1024, # MB
min_containers=1,
region="us-west",
)
@modal.concurrent(target_inputs=64)
@modal.experimental.http_server(port=API_PORT, proxy_regions=["us-west"])
class DeepgramServer(DeepgramServerBase):
...Configure hardware
Section titled “Configure hardware”For Deepgram’s hardware minimums, see Deployment Environments → Engine.
For Modal’s GPU options, see Modal: GPU.
Configure autoscaling
Section titled “Configure autoscaling”Modal automatically scales the number of Deepgram containers up and down based on per-container concurrency.
See their Scaling Out guide and Input Conccurrency guide for the different parameters and their functionality. Note that not all available parameters are surfaced in app.py.
Configure regions
Section titled “Configure regions”To optimize network latency, you will likely want to set the PROXY_REGION AND SERVER_REGION and route traffic from clients in those regions to that deployment.
PROXY_REGIONspecifies the location of the Modal proxy that routes requests to containers. It can take one of four values:us-east,us-west,eu-west,ap-south.SERVER_REGIONspecifies which region(s) the server containers can reside in. See the Modal Region Selection doc for more information.