Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Speed, Pause, Pronunciation

Aura-2 Controls enable fine-grained adjustments to speech output, allowing you to modify speaking speed and override pronunciation for specific words. These controls are designed for enterprise use cases requiring precise voice quality for industry-specific terminology, brand names, and complex content.

Control REST WebSocket Languages
Speed control Yes Yes English (en), Spanish (es)
Pronunciation control Yes Yes English (en), Spanish (es)

Adjust the speaking rate of generated audio. Speed control modifies the pace of speech while maintaining natural prosody and voice quality.

Parameter Location Type Default Range Description
speed query float 1.0 0.7 - 1.5 Speaking rate multiplier
Bash
curl --request POST \
     --header "Content-Type: application/json" \
     --header "Authorization: Token DEEPGRAM_API_KEY" \
     --output your_output_file.mp3 \
     --data '{"text":"Hello, how can I help you today?"}' \
     --url "https://api.deepgram.com/v1/speak?model=aura-2-thalia-en&speed=0.9"
Value Effect Use Case
0.7 30% slower Language learning, accessibility, legal compliance
0.8 20% slower Complex instructions, elderly users
0.9 10% slower Clear explanations, training content
1.0 Normal speed Default conversational pace
1.1 10% faster Efficient notifications
1.2 20% faster Quick alerts, time-sensitive content
1.5 50% faster Rapid playback, content preview

Override the default pronunciation of specific words using International Phonetic Alphabet (IPA) notation.

Pronunciation overrides are specified inline within the text using escaped JSON objects:

\{"word": "dupilumab", "pronounce": "duːˈpɪljuːmæb"\}

Where:

  • word is the original text (used for billing and display)
  • pronounce is the IPA phonetic transcription
  • Curly braces must be escaped with backslashes (\{ and \})
Bash
curl -X POST "https://api.deepgram.com/v1/speak?model=aura-2-thalia-en&speed=0.8" \
     -H "Authorization: token DEEPGRAM_API_KEY" \
     -H "Content-Type: application/json" \
     --output your_output_file.mp3 \
     -d '{"text": "Take \\{\"word\": \"Azathioprine\", \"pronounce\": \"æzəˈθaɪəpriːn\"\\} twice daily with \\{\"word\": \"dupilumab\", \"pronounce\": \"duːˈpɪljuːmæb\"\\}."}'
Category Word IPA Spoken As
Medical dupilumab duːˈpɪljuːmæb ”doo-PIL-yoo-mab”
Medical azathioprine æzəˈθaɪəpriːn ”az-uh-THIGH-oh-preen”
Brand Hermès ɛərˈmɛz ”air-MEZ”
Personal name Nguyen ˈwɪn ”win”
Technical SQL ˈsiːkwəl ”sequel”

A few rules of thumb for producing IPA for your own vocabulary:

Best practices:

  • Always validate by ear. IPA that looks correct on the page can still sound off when synthesized — listen to the output before shipping.
  • Match the dialect. UK and US pronunciations differ (e.g., schedule, aluminum). Make sure the IPA you choose matches the voice and audience you’re targeting.
Rule Limit
Max pronunciations per request 500
Max IPA string length 128 characters
IPA length ratio Cannot exceed 10x the source word length (min floor = 15)
Max input text length 2000 characters

Speed and pronunciation controls can be used together in the same request.

Python
from deepgram import DeepgramClient
from deepgram.core.request_options import RequestOptions

client = DeepgramClient(api_key="YOUR_API_KEY")

# Speed control via request_options
request_opts = RequestOptions(additional_query_parameters={"speed": "0.8"})

# Inline IPA replacements with escaped curly braces
text = r'Take \{"word": "Azathioprine", "pronounce": "æzəˈθaɪəpriːn"\} twice daily with \{"word": "dupilumab", "pronounce": "duːˈpɪljuːmæb"\}.'

response = client.speak.v1.audio.generate(
    text=text,
    model="aura-2-thalia-en",
    encoding="mp3",
    request_options=request_opts
)

audio_bytes = b"".join(response)
with open("medical_instructions.mp3", "wb") as f:
    f.write(audio_bytes)
Python
from deepgram import DeepgramClient

client = DeepgramClient(api_key="YOUR_API_KEY")

# Ensure consistent brand pronunciation with escaped braces
text = 'Visit \\{"word": "Hermès", "pronounce": "ɛərˈmɛz"\\} for the latest collection.'

response = client.speak.v1.audio.generate(
    text=text,
    model="aura-2-thalia-en",
    encoding="mp3"
)

audio_bytes = b"".join(response)
with open("brand_pronunciation.mp3", "wb") as f:
    f.write(audio_bytes)
Symbol Example As in
iː /biːt/ beat
ɪ /bɪt/ bit
eɪ /beɪt/ bait
ɛ /bɛt/ bet
æ /bæt/ bat
ɑː /fɑːðər/ father
ɔː /kɔːt/ caught
oʊ /boʊt/ boat
ʊ /pʊt/ put
uː /buːt/ boot
ʌ /kʌt/ cut
ə /əˈbaʊt/ about
Symbol Example As in
p /pɪn/ pin
b /bɪn/ bin
t /tɪn/ tin
d /dɪn/ din
k /kæt/ cat
ɡ /ɡɛt/ get
f /fɪn/ fin
v /væn/ van
θ /θɪŋk/ think
ð /ðæt/ that
s /sɪt/ sit
z /zɪp/ zip
ʃ /ʃɪp/ ship
ʒ /ˈvɪʒən/ vision
h /hæt/ hat
tʃ /tʃɪp/ chip
dʒ /dʒʌmp/ jump
m /mæn/ man
n /nɛt/ net
ŋ /sɪŋ/ sing
l /lɛt/ let
r /rɛd/ red
w /wɪn/ win
j /jɛs/ yes
Symbol Meaning Example
ˈ Primary stress /ˈæp.əl/ (apple)
ˌ Secondary stress /ˌɪn.fərˈmeɪ.ʃən/ (information)
Control Billing behavior
Speed Not billed - adjusting rate doesn’t affect billing
Pronunciation Billed by underlying word - IPA input is not billed

Example: Hello, \{"word": "Mr.", "pronounce": "ˈmɪstɚ"\} Bond. is billed as Hello, Mr. Bond. (16 characters)

HTTP/1.1 200 OK
content-type: audio/mpeg
dg-request-id: req_xyz789
dg-model-name: aura-2-thalia-en
dg-char-count: 47
dg-pronunciations-applied: 2
dg-speed-used: 0.8
Header Description
dg-pronunciations-applied Number of pronunciation overrides applied
dg-speed-used Effective speaking rate used
dg-pronunciation-warnings Non-fatal warnings for invalid IPA
JSON
{"err_code": "speed_out_of_range", "err_msg": "Speed must be between 0.7 and 1.5"}
JSON
{"err_code": "pronunciation_invalid", "err_msg": "Invalid IPA notation for 'azathioprine'"}
Limit Value
Max input text length 2000 characters
Speed range 0.7 - 1.5
Max pronunciations per request 500
Max IPA string length 128 characters
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu