Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Changelog

Switch Flux STT to digits for a PIN, phone number, or order number, then back to words, without reconnecting. To turn Numerals on or off during a Flux STT stream, send numerals as a boolean in a Configure message:

JSON
{ "type": "Configure", "numerals": true }

The update applies to transcripts Flux STT sends after it processes the message, and ConfigureSuccess now includes numerals in the full active configuration it echoes. The numerals query parameter still sets the initial value when the stream opens. This replaces the connection-time-only behavior in the July 17 entry.

Mid-stream numerals work on flux-general-en and on flux-general-multi for English, Spanish, French, German, Russian, Portuguese, Italian, and Dutch.

Nova-3 Improved Models for Flemish, German (Switzerland), Lithuanian, and Portuguese

Section titled “Nova-3 Improved Models for Flemish, German (Switzerland), Lithuanian, and Portuguese”

We have released improved Nova-3 monolingual models for four existing languages. These updates enhance transcription quality for batch and streaming workloads.

  • German (Switzerland) (de-CH)
  • Portuguese (pt, pt-BR, pt-PT)
  • Lithuanian (lt)
  • Flemish (nl-BE)

You do not need to change your API requests to access the improvements to existing languages; the updates are live for all users.

For more details and to see the full list of supported languages, visit the Models & Languages Overview page.

For self-hosted customers, please request the updated models for download from your Deepgram account representative.

Nova-3 Improved Models for Danish, Estonian, Flemish, Italian, Lithuanian, Macedonian, Polish, Urdu, and Vietnamese

Section titled “Nova-3 Improved Models for Danish, Estonian, Flemish, Italian, Lithuanian, Macedonian, Polish, Urdu, and Vietnamese”

We have released improved Nova-3 monolingual models for nine existing languages. These updates enhance transcription quality for batch and streaming workloads.

  • Polish (pl)
  • Danish (da)
  • Vietnamese (vi)
  • Flemish (nl-BE)
  • Estonian (et)
  • Macedonian (mk)
  • Urdu (ur)
  • Italian (it)
  • Lithuanian (lt)

You do not need to change your API requests to access the improvements to existing languages; the updates are live for all users.

For more details and to see the full list of supported languages, visit the Models & Languages Overview page.

For self-hosted customers, please request the updated models for download from your Deepgram account representative.

Correction: Browser Agent SDK has no client-side VAD

Section titled “Correction: Browser Agent SDK has no client-side VAD”

The Browser Agent SDK never shipped client-side Silero voice activity detection. Documentation published with the SDK described it as an available feature, and that documentation was wrong. The following did not exist in any released package:

  • The vad option on MicrophoneOptions in @deepgram/agents.
  • The speechThreshold and silenceThreshold VAD options.
  • The speech-start and speech-end events on AgentMicrophone.
  • The vad field in WidgetConfig and the widget's VAD configuration section.
  • The @ricky0123/vad-web and onnxruntime-web peer dependencies. Uninstall them if you added them for VAD.

MicrophoneOptions accepts sampleRate, echoCancellation, noiseSuppression, and autoGainControl. The microphone streams continuously while it is unmuted.

Barge-in works, and it is driven server-side. The Voice Agent API emits user-started-speaking, which you pair with player.interrupt() to stop playback:

JavaScript
session.on("user-started-speaking", () => player.interrupt());

@deepgram/react subscribes to that event and interrupts playback for you, so applications built on the provider and hooks need no change.

The Browser Agent pages now reflect the shipped API. See JavaScript SDK, Widget Embedding Guide, and React Hooks & Provider.

The server-side token-minting example on the Browser Agent SDK overview posted {"ttl": 30} to POST /v1/auth/grant. The field is named ttl_seconds:

JavaScript
body: JSON.stringify({ ttl_seconds: 30 }),

/v1/auth/grant accepts unrecognized fields, returns 200, and issues a token with the 30-second default, so a request built from the old example raised no error while ignoring the requested lifetime. If you copied that example and increased the number, your tokens expired after 30 seconds.

A 30-second token is enough for a full call. The token only has to be valid at the WebSocket handshake, and the connection stays open afterward. Minting a token requires an API key with Member permissions or higher. See Token-Based Authentication.

Nova-3 Pharma: Speech-to-Text for Pharmaceutical Use Cases (English)

Section titled “Nova-3 Pharma: Speech-to-Text for Pharmaceutical Use Cases (English)”

We have released nova-3-pharma, a new Nova-3 model purpose-built for pharmaceutical vocabulary, with a focus on accurate drug-name recognition. It is designed for pharmacy and healthcare voice-agent workflows where getting the medication right matters most. Available in English for both batch and streaming.

Nova-3 Pharma is now available through our API. To access:

  • Use model=nova-3-pharma in your API calls
  • Available for hosted customers
    • For self-hosted customers, please request the model from your Deepgram account representative
  • Supports both pre-recorded (batch) and real-time streaming transcription
  • Supported languages: English (en, en-US, en-AU, en-CA, en-GB, en-IE, en-IN, en-NZ)

For more details, visit the Models & Languages Overview and Model pages.

The Deepgram India endpoint (api.in.deepgram.com) is now generally available for Speech-to-Text, Text-to-Speech, Voice Agent, and Text Intelligence APIs.

The Deepgram India endpoint (api.in.deepgram.com) is now generally available for customers requiring data processing within India.

The India endpoint supports the following Deepgram APIs:

  • Speech-to-Text: /v1/listen and /v2/listen
  • Text-to-Speech: /v1/speak and /v2/speak
  • Voice Agent: /v1/agent/converse
  • Text Intelligence: /v1/read

To use the India endpoint, replace api.deepgram.com with api.in.deepgram.com in your SDK or API requests. Your existing API keys and tokens work with the India endpoint.

For WebSocket connections, use the corresponding India URLs:

  • Speech-to-Text: wss://api.in.deepgram.com/v1/listen
  • Speech-to-Text (Flux): wss://api.in.deepgram.com/v2/listen
  • Text-to-Speech: wss://api.in.deepgram.com/v1/speak
  • Text-to-Speech (Flux): wss://api.in.deepgram.com/v2/speak
  • Voice Agent: wss://api.in.deepgram.com/v1/agent/converse

For detailed configuration instructions and SDK examples, see our Regional Endpoints documentation. For what Deepgram guarantees about where your data is processed and stored, see Your Data at Deepgram.

Deepgram Self-Hosted release 260915 adds general availability support for NVIDIA Blackwell GPUs, forced end-of-turn and mid-stream numeral configuration for Flux, and improved German formatting.

Deepgram Self-Hosted September 2026 Release (260915)

Section titled “Deepgram Self-Hosted September 2026 Release (260915)”
  • quay.io/deepgram/self-hosted-api:release-260915

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-api:1.192.37
  • quay.io/deepgram/self-hosted-engine:release-260915

    • Equivalent image to:

      • quay.io/deepgram/self-hosted-engine:3.131.2
    • Minimum required NVIDIA driver version: >=580, open kernel module flavor

      • The minimum driver version has increased this release, up from >=570.172.08. The Engine image is built against CUDA 13, and NVIDIA's official CUDA 13 support begins with the 580 driver branch. NVIDIA's documentation states the CUDA 13 minimum as >=580 rather than a specific version.
      • The driver must be the open kernel module flavor, for example nvidia-driver-580-open. The proprietary build of the same version is not supported.
        • On Blackwell hardware the consequence is total: NVIDIA never added Blackwell support to the proprietary 580 branch, so on that build a Blackwell GPU disappears from the system, nvidia-smi reports no devices, and Engine will not start. See Drivers and Containerization Platforms for installation and verification steps.
  • quay.io/deepgram/self-hosted-license-proxy:release-260915

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-license-proxy:1.11.1
  • quay.io/deepgram/self-hosted-billing:release-260915

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-billing:1.14.0-2

FIPS 140-3 images are available for all four self-hosted services, published under the -fips tag suffix:

  • quay.io/deepgram/self-hosted-api:release-260915-fips

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-api:1.192.37-fips
  • quay.io/deepgram/self-hosted-engine:release-260915-fips

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-engine:3.131.2-fips
  • quay.io/deepgram/self-hosted-license-proxy:release-260915-fips

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-license-proxy:1.11.1-fips
  • quay.io/deepgram/self-hosted-billing:release-260915-fips

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-billing:1.14.0-2-fips

These images run OpenSSL's FIPS 140-3 validated cryptographic provider. Enable FIPS by setting [fips] mode = "enabled" in each service's configuration. FIPS mode should be enabled only on the -fips images; the standard images do not include the FIPS provider. See FIPS-Compliant Deployment for setup and configuration details.

  • NVIDIA Blackwell GPU Support (GA) — Blackwell-generation GPUs are now generally available for self-hosted deployments. See Drivers and Containerization Platforms for driver requirements and Model and GPU Compatibility for which models run on which GPUs.
  • Flux Forced End-of-Turn — your client can now end a Flux turn on demand rather than waiting for the model to detect the end of speech. Enable it by setting listen_v2 = true and listen_v2_force_end_turn = true under [features] in your API configuration, then send a ForceEndTurn message on an active /v2/listen stream. The resulting EndOfTurn event reports trigger: "manual", and its transcript covers only the audio received before the message. See Force End Turn for details.
  • Mid-Stream Numeral Configuration for Flux — change numeral formatting during an active /v2/listen stream by sending numerals in a Configure control message, instead of fixing it at connection time. Previously numerals could only be set as a query parameter when the stream opened.
  • Formatting Improvements — improved reading of German numbers, currency, times, and addresses.
  • General Improvements — dependency and security updates, and refreshed base images.

Deepgram Python SDK v7.9.0, JavaScript SDK v5.11.0, and Java SDK v0.10.0 are now available. All three releases add Voice Agent ForceEndTurn support, Flux TTS expressivity controls for Agent providers, and the expanded Flux TTS speed range of 0.5 through 1.5 in 0.05 increments.

Deepgram Python SDK v7.9.0 adds send_force_end_turn() and typed ForceEndTurn messages for Voice Agent sessions using a Deepgram V2 (Flux STT) listen provider. It also adds integer Agent Flux TTS expressivity from -2 through 2 and the expanded Flux TTS speed range. The unsupported aura-2-perseo-it model is removed from Speak V1 model literals; the API never served that model.

For release details, see deepgram-python-sdk v7.9.0.

Deepgram JavaScript SDK v5.11.0 adds socket.sendForceEndTurn() and Agent Flux TTS expressivity from -2 through 2. The former SpeakV2Speed constants and Aura2PerseoIt model constants remain as deprecated aliases for source compatibility.

For release details, see deepgram-js-sdk v5.11.0.

Deepgram Java SDK v0.10.0 adds sendForceEndTurn(...), Agent Flux TTS expressivity, and the expanded Flux TTS speed range.

This pre-1.0 release includes breaking changes: it removes SpeakV2Speed and its named speed constants, plus the unsupported AURA2PERSEO_IT model constants and visitor methods. Configure speed directly with a numeric value, such as .speed(1.05), and use a supported Speak V1 model.

For release details, see deepgram-java-sdk v0.10.0. For migration steps, see the v0.9 to v0.10 migration guide.

@deepgram/react 0.2.0 adds runtime Voice Agent controls and makes React session lifecycle handling safer. The release updates the React layer for @deepgram/agents 0.1.2 and @deepgram/sdk 5.9.0.

Use sendAgentMessage() to inject an agent message with "default", "queue", or "interrupt" behavior. Use updateListen(), updateThink(), updateSpeak(), and updatePrompt() to change supported agent settings during a session.

The provider and standalone hook now expose typed callbacks for protocol notifications, including confirmations for each runtime settings update. useAgentMode() and useDeepgramAgent() also expose the "thinking" mode and isThinking state.

Start, stop, reconnect, and React StrictMode handling now clean up canceled or failed microphone, playback, and client-tool resources. Automatic start failures call onSdkError, or console.error when no callback is set. A manual start() call rejects instead.

This is a 0.2.0 release. Update exhaustive AgentMode switches to handle "thinking"; public context and hook result types add required members; registerClientTool() returns an unsubscribe function; and useDeepgramAgent().start() starts a fresh session and clears conversation.

See React Hooks & Provider and the migration guide for API details and upgrade examples.

Hold a function call until the user's turn is confirmed

Section titled “Hold a function call until the user's turn is confirmed”

The agent begins building a reply before speech-to-text confirms the user has finished speaking. Agent audio is held until that confirmation, but function calls have always dispatched as soon as the LLM emitted them. For a function with an irreversible side effect, such as ending a call or booking an appointment, that meant the action could run for a turn the user then continued.

Add defer_until_eot: true to any function in agent.think.functions to hold it until the turn is confirmed. If the user keeps speaking and the turn resumes, the call is discarded before it does anything.

JSON
{
  "name": "end_call",
  "description": "End the conversation and close the connection",
  "parameters": { "type": "object", "properties": {} },
  "defer_until_eot": true
}

The default is false, so existing agents are unchanged. Deferring one function does not delay the others: a function that did not opt in still dispatches immediately, even with a deferred call ahead of it in the same turn.

defer_until_eot applies to every listen provider, not only Flux.

Clients now receive an explicit signal when a client-side function call they were sent is cancelled because the user started speaking again, either inside the speculative window (the turn resumed) or after the turn was confirmed (barge-in). Only calls the client already received are announced.

JSON
{
  "type": "FunctionCallCancelled",
  "functions": [{ "id": "fc_12345678-90ab-cdef-1234-567890abcdef", "name": "book_appointment" }]
}

Stop work on that id and send no FunctionCallResponse for it.

For details, see Speculative Replies & Turn Confirmation, Function Call Cancelled, and Configure the Voice Agent.

Nova-3 Adds Kazakh, Plus Improved Models for Estonian, Hebrew, Latvian, Lithuanian, Macedonian, Malay, and Polish

Section titled “Nova-3 Adds Kazakh, Plus Improved Models for Estonian, Hebrew, Latvian, Lithuanian, Macedonian, Malay, and Polish”

We have added Kazakh as a new Nova-3 language and released improved Nova-3 monolingual models for seven existing languages. These updates expand language coverage and enhance transcription quality for batch and streaming workloads.

  • Kazakh (kk, kk-KZ) — available for batch and streaming.

Access this language by setting model="nova-3" and language="kk" in your request.

  • Polish (pl)
  • Hebrew (he)
  • Macedonian (mk)
  • Estonian (et)
  • Latvian (lv)
  • Lithuanian (lt)
  • Malay (ms)

You do not need to change your API requests to access the improvements to existing languages; the updates are live for all users.

For more details and to see the full list of supported languages, visit the Models & Languages Overview page.

For self-hosted customers, please request the updated models for download from your Deepgram account representative.

Flux TTS speed now runs from 0.5 to 1.5 in 0.05 increments, up from 0.85–1.15. The default stays 1.0, and every previously accepted value still works.

The wider range applies everywhere speed is accepted on Flux TTS: the /v2/speak streaming WebSocket (at connection and mid-session with Configure), the /v2/speak batch REST endpoint, and agent.speak.provider.speed in Voice Agent sessions running provider.version v2. A value outside the range is still rejected with SPEED_OUT_OF_RANGE, and one off the 0.05 increment with SPEED_INCREMENT_INVALID.

Aura speed is unchanged at 0.7–1.5.

For details, see Getting Started with Flux TTS, Client Messages, and TTS Models.

Flux now gives you full control over how turns end. Beyond Flux's native end-of-turn detection, you can override it, suppress it entirely, or blend the two — and switch between these modes on the fly, even for a single turn, with a Configure message. Turn-taking goes from a fixed behavior to something you shape around each moment of a conversation.

  • Automatic (default). Flux's native end-of-turn detection decides when a turn ends. EndOfTurn events carry "trigger": "model". Nothing changes for existing integrations.
  • Semi-manual. Keep Flux's detection running, but override it whenever you have a definitive signal — a push-to-talk release, a DTMF tone, a "send" tap — by sending a ForceEndTurn. The model handles the ambiguous endings; you handle the unambiguous ones.
  • Fully manual. Set eot_threshold=1.0 to suppress native detection entirely and own every turn boundary yourself, driving each ending with ForceEndTurn. Ideal when you already run your own VAD or endpointing stack. See Bring Your Own Turn Detection.

Because eot_threshold can be updated live with a Configure message, you can change turn-taking behavior for exactly the turn that needs it and then change back. For example, when a caller is about to read an 8-character alphanumeric account ID, raise eot_threshold to 1.0 so Flux won't cut them off between characters, end the turn yourself once you've collected all 8, then drop back to the default for natural conversation:

JSON
{ "type": "Configure", "thresholds": { "eot_threshold": 1.0 } }

Flux replies with ConfigureSuccess and the new behavior applies immediately, without reconnecting.

ForceEndTurn message. Send {"type": "ForceEndTurn"} as a WebSocket text frame to end the active turn immediately. Flux ends the turn on the audio transcribed so far and emits a standard EndOfTurn. See Force End Turn.

trigger field on EndOfTurn. Every EndOfTurn event now includes a trigger field stating what ended the turn:

JSON
{
  "event": "EndOfTurn",
  "transcript": "I need to cancel my subscription",
  "end_of_turn_confidence": 0.35,
  "trigger": "manual"
}
  • model — Flux's native end-of-turn detection.
  • manual — a ForceEndTurn message.
  • timeout — the eot_timeout_ms safety net.

eot_threshold maximum raised to 1.0. Set eot_threshold=1.0 — at connect time or mid-stream via Configure — to fully suppress natural end-of-turn detection.

These changes are additive and backward-compatible. Clients that don't send ForceEndTurn see no behavior change beyond the new trigger field on EndOfTurn events. Flux ForceEndTurn and Configure-based turn control are supported on the /v2/listen endpoint, and are available on the global endpoint as well as Deepgram's EU (api.eu.deepgram.com) and AU (api.au.deepgram.com) regional endpoints.

The Voice Agent API provides the same turn-taking control through the ForceEndTurn client message when the agent uses the Flux (v2) listen provider. Send it to end the current user turn immediately, or set agent.listen.provider.eot_threshold=1.0 to fully suppress natural end-of-turn detection and drive every turn ending with ForceEndTurn.

Nova-3 Model Improvements for Bulgarian, Croatian, Estonian, Georgian, Italian, Latvian, Lithuanian, Malay, Marathi, and Telugu

Section titled “Nova-3 Model Improvements for Bulgarian, Croatian, Estonian, Georgian, Italian, Latvian, Lithuanian, Malay, Marathi, and Telugu”

We have released improved Nova-3 monolingual models for ten existing languages. These updates enhance transcription quality for batch and streaming workloads.

  • Bulgarian (bg)
  • Croatian (hr)
  • Estonian (et)
  • Italian (it)
  • Latvian (lv)
  • Lithuanian (lt)
  • Malay (ms)
  • Georgian (ka, ka-GE)
  • Marathi (mr)
  • Telugu (te)

You do not need to change your API requests to access these improvements; the updates are live for all users.

For more details and to see the full list of supported languages, visit the Models & Languages Overview page.

For self-hosted customers, please request the updated models for download from your Deepgram account representative.

Nova-3 Adds Assamese, Mongolian, and Pashto, Plus Improved Models for Czech, Danish, Swedish, Tagalog, and Turkish

Section titled “Nova-3 Adds Assamese, Mongolian, and Pashto, Plus Improved Models for Czech, Danish, Swedish, Tagalog, and Turkish”

We have added Assamese, Mongolian, and Pashto as new Nova-3 languages and released improved Nova-3 monolingual models for several existing languages. These updates expand language coverage and enhance transcription quality for batch and streaming workloads.

  • Assamese (as, as-IN) — available for batch and streaming.
  • Mongolian (mn) — available for batch and streaming.
  • Pashto (ps, ps-AF) — available for batch and streaming.

Access these languages by setting model="nova-3" and the relevant language code in your request.

  • Czech (cs, cs-CZ)
  • Danish (da, da-DK)
  • Tagalog (tl)
  • Turkish (tr, tr-TR)
  • Swedish (sv, sv-SE)

You do not need to change your API requests to access the improvements to existing languages; the updates are live for all users.

For more details and to see the full list of supported languages, visit the Models & Languages Overview page.

For self-hosted customers, please request the updated models for download from your Deepgram account representative.

Deepgram Self-Hosted release 260826 adds expressivity control to Flux TTS, and improves currency, date, and Japanese punctuation formatting.

Deepgram Self-Hosted August 2026 Release (260826)

Section titled “Deepgram Self-Hosted August 2026 Release (260826)”
  • quay.io/deepgram/self-hosted-api:release-260826

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-api:1.192.31
  • quay.io/deepgram/self-hosted-engine:release-260826

    • Equivalent image to:

      • quay.io/deepgram/self-hosted-engine:3.128.0
    • Minimum required NVIDIA driver version: >=570.172.08

  • quay.io/deepgram/self-hosted-license-proxy:release-260826

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-license-proxy:1.11.0-1
  • quay.io/deepgram/self-hosted-billing:release-260826

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-billing:1.14.0-1

FIPS 140-3 images are available for all four self-hosted services, published under the -fips tag suffix:

  • quay.io/deepgram/self-hosted-api:release-260826-fips

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-api:1.192.31-fips
  • quay.io/deepgram/self-hosted-engine:release-260826-fips

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-engine:3.128.0-fips
  • quay.io/deepgram/self-hosted-license-proxy:release-260826-fips

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-license-proxy:1.11.0-1-fips
  • quay.io/deepgram/self-hosted-billing:release-260826-fips

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-billing:1.14.0-1-fips

These images run OpenSSL's FIPS 140-3 validated cryptographic provider. Enable FIPS by setting [fips] mode = "enabled" in each service's configuration. FIPS mode should be enabled only on the -fips images; the standard images do not include the FIPS provider. See FIPS-Compliant Deployment for setup and configuration details.

  • Flux TTS General Availability — Flux TTS is now generally available (GA) on self-hosted.
  • Flux TTS Expressivity — a new expressivity control makes a Flux TTS voice more calm or more animated. Ask your Deepgram representative for the accompanying new model version f94b1bd5.
  • Wider Flux TTS speed Range — Flux TTS speed now runs from 0.5 to 1.5 in 0.05 increments, up from 0.85–1.15.
  • Smart Formatting Improvements — improved currency and date formatting, and improved Japanese end-of-sentence punctuation.
  • General Improvements — dependency and security updates, and refreshed base images.

You can now end the current user turn from your own signal instead of waiting for end-of-turn detection.

Send ForceEndTurn to end the turn in progress immediately. The agent treats everything transcribed so far as the complete user utterance and moves on to think and speak.

JSON
{
  "type": "ForceEndTurn"
}

Use it for push-to-talk release, DTMF tones, UI actions, or to drive turn taking from your own VAD or endpointing stack.

ForceEndTurn requires a Deepgram V2 (Flux) listen provider (agent.listen.provider.version set to v2). With any other listen provider the server sends a FORCE_END_TURN_UNSUPPORTED warning and the turn does not end.

For details, see Force End Turn.

agent.listen.provider.eot_threshold now accepts values up to 1.0. Set it to 1.0 to fully suppress natural end-of-turn detection and own turn taking with ForceEndTurn.

For details, see Configure the Voice Agent and Bring Your Own Turn Detection.

Deepgram CLI 0.3.0 is now available. It completes CLI support for Flux TTS, which reached general availability on 12 August, fixes Flux STT streaming, and enforces the CLI's exit-code contract.

Upgrading from 0.2.26 picks up everything in this release. 0.2.27 was tagged but never published, so its changes arrive here.

dg speak now defaults to flux-alexis-en on Speak v2 (WebSocket streaming) instead of Aura 2. Synthesised audio differs from earlier releases unless you request an Aura model explicitly:

Shell
dg speak "Hello from Deepgram" -m aura-2-asteria-en -o hello.mp3

Two Flux-only controls are now available. Both are validated before the request and rejected for non-flux-* models.

Option Values Notes
--speed 0.85 to 1.15 in 0.05 steps 1.00 is nominal
--expressivity -2 to 2 Beta. 0 is nominal; negative is flatter, positive more animated
Shell
dg speak "So exciting!" --expressivity 2 --speed 1.05 -o lively.wav

Any GA voice works with -m. For the full catalog see Voices & Languages; dg models lists Aura voices only.

Flux streaming returns raw audio: linear16 by default, plus mulaw and alaw. The --container option and the wider --encoding set apply to Aura only. Aura is otherwise unchanged, Spanish voices included.

dg speak synthesises one turn and exits, so the GA barge-in surface (Interrupt and SpeechInterrupted) and mid-session Configure are intentionally not exposed. Those target live agent pipelines; --speed and --expressivity are fixed for the connection here. See Interruption Handling if you need them.

dg listen has routed flux-* models to Listen v2 since 0.2.5, but v2 was receiving v1-only parameters (language, smart_format, punctuate, channels, diarize, interim_results) and returning HTTP 400. Flux STT streaming now works:

Shell
dg listen --mic --model flux-general-en

Also fixed in this release:

  • TurnInfo events are assembled per turn, and finite streams flush the final in-flight turn instead of dropping it
  • Fatal error frames propagate their code and description, and exit non-zero
  • Multichannel raw audio is rejected before streaming rather than sent as stereo bytes read as mono
  • Partial and null word timings are backfilled, so saved captions are valid
  • --diarize warns and is cleared consistently, since Flux does not support it

Flux STT is streaming only. Files and URLs use Listen v1.

Shell
dg listen call.wav --redact numbers --numerals
  • --redact is repeatable. Flux STT accepts numbers and aggressive_numbers; Listen v1 also accepts categories such as pci and ssn.
  • --numerals converts spoken numbers to digits.

Both apply to prerecorded and live paths. All of the above runs on Deepgram Python SDK 7.7.0.

This release includes a breaking change. dg previously exited 0 regardless of outcome. It now follows the documented contract:

Code Meaning
0 Success
1 Error, including crashes and usage errors such as an unknown command or invalid flag
2 User interrupt: Ctrl-C, or Ctrl-D at a prompt

Scripts and CI steps that ignored the exit code will now surface failures they were previously swallowing. No command that succeeds changes its exit code. See Exit Codes.

Usage errors, crashes, and cancellation messages are written to stderr instead of stdout, so dg -o json keeps stdout machine-readable when a command fails this way. Previously these printed to stdout, which corrupted output piped into tools such as jq.

This covers the root command path. Some command-level errors, an authentication failure among them, still print to stdout, so scripts should branch on the exit code rather than assume stdout parses.

The CLI ships as a root package plus per-command packages. Root's dependency floors were lower than the versions being published, so pip's default only-if-needed strategy left most command packages at their installed version: dg --version reported the new release while its fixes had not arrived. This is why Flux TTS, published in 0.2.26, never reached anyone who upgraded rather than installing fresh.

Floors now match the published versions, so both of these deliver the full release:

Shell
dg update
pip install --upgrade deepctl

Installations managed with uv, Homebrew, or the install script were not affected.

  • -o yaml and -o csv no longer drop square-bracketed text from values
  • dg keys --delete KEY_ID asks for confirmation instead of reporting Cancelled by user without deleting
  • dg keys --create --dry-run reports what it would create instead of failing
  • dg mcp handles a closed stdio pipe during startup notifications and on the error path

For release details, see deepgram/cli v0.3.0.

Expressivity for Flux TTS voices, and two new OpenAI models

Section titled “Expressivity for Flux TTS voices, and two new OpenAI models”

agent.speak.provider.expressivity shifts a Flux TTS voice's delivery register along a calm to animated axis. It accepts the whole numbers -2 to 2 and defaults to 0, the voice's tuned delivery. Negative values produce calmer, steadier delivery; positive values produce more animated delivery with a wider pitch range. Every Flux voice supports it, and the value applies for the whole session.

JSON
{
  "agent": {
    "speak": {
      "provider": {
        "type": "deepgram",
        "version": "v2",
        "model": "flux-haley-en",
        "expressivity": -1
      }
    }
  }
}

expressivity is a beta parameter: its behavior may be tuned in future model versions, and non-default values raise the chance of hallucinations and pronunciation errors, so audition the value you plan to ship. 0 remains the only value validated for production.

For value-by-value guidance, see Expressivity and Configure the Voice Agent.

Deepgram's managed OpenAI provider adds two models:

Model Pricing Tier
gpt-5.6-luna Standard
gpt-5.6-terra Advanced
JSON
{
  "agent": {
    "think": {
      "provider": {
        "type": "open_ai",
        "model": "gpt-5.6-luna"
      }
    }
  }
}

For the full catalog, see LLM Models.

Nova-3 Adds Afrikaans and Georgian, Plus Improved Models for Hungarian, Macedonian, Russian, Slovak, Slovenian, and Urdu

Section titled “Nova-3 Adds Afrikaans and Georgian, Plus Improved Models for Hungarian, Macedonian, Russian, Slovak, Slovenian, and Urdu”

We have added Afrikaans and Georgian as new Nova-3 languages and released improved Nova-3 monolingual models for several existing languages. These updates expand language coverage and enhance transcription quality for batch and streaming workloads.

  • Afrikaans (af, af-ZA) — available for batch and streaming.
  • Georgian (ka, ka-GE) — available for batch and streaming.

Access these languages by setting model="nova-3" and the relevant language code in your request.

  • Russian (ru)
  • Slovak (sk)
  • Slovenian (sl)
  • Urdu (ur)
  • Hungarian (hu)
  • Macedonian (mk)

You do not need to change your API requests to access the improvements to existing languages; the updates are live for all users.

For more details and to see the full list of supported languages, visit the Models & Languages Overview page.

For self-hosted customers, please request the updated models for download from your Deepgram account representative.

Numerals Support Now Available for 4 New Languages: Bulgarian, Chinese (Cantonese, Traditional), Malay, and Korean (Monolingual Models)

Section titled “Numerals Support Now Available for 4 New Languages: Bulgarian, Chinese (Cantonese, Traditional), Malay, and Korean (Monolingual Models)”
  • Bulgarian (bg)
  • Chinese (Cantonese, Traditional) (zh-HK)
  • Malay (ms)
  • Korean (ko, ko-KR)

You can now use Deepgram’s Numerals feature with monolingual models for Bulgarian, Chinese (Cantonese, Traditional), Malay, and Korean. Numerals converts spoken numbers into digits (for example, "three hundred" → "300") in your transcript, helping you create more accurate and easily processed results.

How to use Numerals:
To enable numerals, add the numerals=true parameter to your Deepgram API request.

Learn more about using Numerals and see the full list of supported languages on the Numerals documentation page.

Showing the 20 most recent of 224 entries. Append /llms.txt to the changelog URL for the complete index.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu