LLM Models
Defines the LLM (Large Language Model) to be used with your Agent. The provider.type field specifies the format or protocol of the API.
For example:
open_aimeans the API follows OpenAI's Chat Completions format.- This option can be used with OpenAI, Azure OpenAI, or Amazon Bedrock — as long as the endpoint behaves like OpenAI's Chat Completion API.
Supported LLM providers
Section titled “Supported LLM providers”| Parameter | open_ai |
anthropic |
aws_bedrock |
google |
groq |
nvidia |
|---|---|---|---|---|---|---|
agent.think.provider.type |
open_ai |
anthropic |
aws_bedrock |
google |
groq |
nvidia |
agent.think.endpoint |
optional | optional | required | optional | required | optional |
The agent.think.endpoint is optional or required based on the provider type:
- For
open_ai,anthropic,google, andnvidia, theendpointfield is optional because Deepgram provides managed LLMs for these providers. - For
groqandaws_bedrockprovider types,endpointis required because Deepgram does not manage those LLMs. - If an
endpointis provided theurlis required butheadersare optional.
See the Amazon Bedrock section below for credentials and endpoint configuration. To fetch the current list of providers and models programmatically, see Listing supported models via the API.
Supported LLM models
Section titled “Supported LLM models”OpenAI
Section titled “OpenAI”| Provider | Model | Pricing Tier |
|---|---|---|
open_ai |
gpt-5.6-terra |
Advanced |
open_ai |
gpt-5.6-luna |
Standard |
open_ai |
gpt-5.5 |
Advanced |
open_ai |
gpt-5.4-nano |
Standard |
open_ai |
gpt-5.4-mini |
Standard |
open_ai |
gpt-5.4 |
Advanced |
open_ai |
gpt-5.3-chat-latest |
Advanced |
open_ai |
gpt-5.2-chat-latest |
Advanced |
open_ai |
gpt-5.2 |
Advanced |
open_ai |
gpt-5.1-chat-latest |
Advanced |
open_ai |
gpt-5.1 |
Advanced |
open_ai |
gpt-5-nano |
Standard |
open_ai |
gpt-5-mini |
Standard |
open_ai |
gpt-5 |
Advanced |
open_ai |
gpt-4.1-nano |
Standard |
open_ai |
gpt-4.1-mini |
Standard |
open_ai |
gpt-4.1 |
Advanced |
open_ai |
gpt-4o-mini |
Standard |
open_ai |
gpt-4o |
Advanced |
Example using Deepgram's managed OpenAI LLM
Section titled “Example using Deepgram's managed OpenAI LLM”JSON
// ... other settings ...
"think": {
"provider": {
"type": "open_ai",
"model": "gpt-4o-mini",
"temperature": 0.5
}
}
// ... other settings ...Example using a custom OpenAI endpoint (BYO)
Section titled “Example using a custom OpenAI endpoint (BYO)”JSON
// ... other settings ...
"think": {
"provider": {
"type": "open_ai",
"model": "gpt-4o-mini",
"temperature": 0.5
},
"endpoint": {
"url": "https://api.openai.com/v1/chat/completions",
"headers": {
"authorization": "Bearer xxxxxxxxx"
}
}
}
// ... other settings ...Anthropic
Section titled “Anthropic”| Provider | Model | Pricing Tier |
|---|---|---|
anthropic |
claude-sonnet-5 |
Advanced |
anthropic |
claude-sonnet-4-6 |
Advanced |
anthropic |
claude-sonnet-4-5 |
Advanced |
anthropic |
claude-haiku-4-5 |
Standard |
anthropic |
claude-3-5-haiku-latest |
Standard |
anthropic |
claude-sonnet-4-20250514 |
Advanced (Deprecated) |
Example using Deepgram's managed Anthropic LLM
Section titled “Example using Deepgram's managed Anthropic LLM”JSON
// ... other settings ...
"think": {
"provider": {
"type": "anthropic",
"model": "claude-haiku-4-5",
"temperature": 0.5
}
}
// ... other settings ...Example using a custom Anthropic endpoint (BYO)
Section titled “Example using a custom Anthropic endpoint (BYO)”JSON
// ... other settings ...
"think": {
"provider": {
"type": "anthropic",
"model": "claude-haiku-4-5",
"temperature": 0.5
},
"endpoint": {
"url": "https://api.anthropic.com/v1/messages",
"headers": {
"x-api-key": "xxxxxxxxx",
"anthropic-version": "2023-06-01"
}
}
}
// ... other settings ...| Provider | Model | Pricing Tier |
|---|---|---|
google |
gemini-3.5-flash |
Standard |
google |
gemini-3.1-flash-lite |
Standard |
google |
gemini-3.1-flash-lite-preview |
Standard (Deprecated May 26, 2025) |
google |
gemini-3-flash-preview |
Standard |
google |
gemini-3-pro-preview |
Advanced |
google |
gemini-2.5-flash |
Standard |
google |
gemini-2.0-flash-lite |
Standard |
Deepgram supports both of Google's APIs for managed Gemini models: the AI Studio API and the Gemini Enterprise Agent (GEA) API, formerly the Vertex API. GEA offers regional endpoints, so requests can stay in the region you select, while global AI Studio offers the broadest model coverage. You can specify either API from any Deepgram Voice Agent endpoint using the think.provider.version parameter; the default depends on which Deepgram endpoint you connect to.
think.provider.version |
Google API |
|---|---|
ai-studio-v1beta |
AI Studio |
gemini-enterprise-agent-v1 |
Gemini Enterprise Agent (GEA) |
v1beta |
AI Studio — alias for ai-studio-v1beta |
For defaults, regional availability, and configuration details, see Regional Endpoints.
Example using Deepgram's managed Google LLM
Section titled “Example using Deepgram's managed Google LLM”JSON
// ... other settings ...
"think": {
"provider": {
"type": "google",
"version": "ai-studio-v1beta",
"model": "gemini-2.5-flash",
"temperature": 0.5
}
}
// ... other settings ...Example using a custom Google endpoint (BYO)
Section titled “Example using a custom Google endpoint (BYO)”When using a custom endpoint, the model property is not supported.
The desired model is specified as part of the endpoint URL instead.
JSON
// ... other settings ...
"think": {
"provider": {
"type": "google",
"temperature": 0.5
},
"endpoint": {
"url": "https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:streamGenerateContent?alt=sse",
"headers": {
"x-goog-api-key": "xxxxxxxxx"
}
}
}
// ... other settings ...NVIDIA
Section titled “NVIDIA”| Provider | Model | Pricing Tier |
|---|---|---|
nvidia |
nemotron-3-nano-30B-A3B |
Standard |
Example using Deepgram's managed NVIDIA LLM
Section titled “Example using Deepgram's managed NVIDIA LLM”JSON
// ... other settings ...
"think": {
"provider": {
"type": "nvidia",
"model": "nemotron-3-nano-30B-A3B",
"temperature": 0.5
}
}
// ... other settings ...| Provider | Model | Pricing Tier |
|---|---|---|
groq |
openai/gpt-oss-20b |
Standard |
Amazon Bedrock
Section titled “Amazon Bedrock”Amazon Bedrock is a BYO provider. Deepgram does not host Bedrock models, so endpoint.url is required and you supply your own AWS credentials. Bedrock model IDs (for example us.anthropic.claude-3-5-sonnet-20241022-v2:0) are passed through to Bedrock as-is.
| Parameter | Value |
|---|---|
agent.think.provider.type |
aws_bedrock |
agent.think.provider.model |
A Bedrock model ID |
agent.think.provider.credentials |
IAM or STS credentials (see below) |
agent.think.endpoint.url |
https://bedrock-runtime.{region}.amazonaws.com/ |
IAM credentials
Section titled “IAM credentials”Use long-lived IAM access keys when your application has stable credentials.
JSON
// ... other settings ...
"think": {
"provider": {
"type": "aws_bedrock",
"model": "us.anthropic.claude-3-5-sonnet-20241022-v2:0",
"temperature": 0.7,
"credentials": {
"type": "iam",
"region": "us-east-2",
"access_key_id": "{{your_access_key_id}}",
"secret_access_key": "{{your_secret_access_key}}"
}
},
"endpoint": {
"url": "https://bedrock-runtime.us-east-2.amazonaws.com/"
}
}
// ... other settings ...STS (temporary) credentials
Section titled “STS (temporary) credentials”Use STS credentials when your application assumes a role and rotates tokens. Add the session_token returned by your STS call.
JSON
// ... other settings ...
"think": {
"provider": {
"type": "aws_bedrock",
"model": "us.anthropic.claude-3-5-sonnet-20241022-v2:0",
"temperature": 0.7,
"credentials": {
"type": "sts",
"region": "us-east-2",
"access_key_id": "{{your_temporary_access_key_id}}",
"secret_access_key": "{{your_temporary_secret_access_key}}",
"session_token": "{{your_session_token}}"
}
},
"endpoint": {
"url": "https://bedrock-runtime.us-east-2.amazonaws.com/"
}
}
// ... other settings ...If you need an OpenAI-compatible proxy in front of Bedrock (for logging, header rewriting, or use of the Bedrock Agents service), see Passing a custom (BYO) LLM through a Cloud Provider below.
Example Payload
Section titled “Example Payload”JSON
// ... other settings ...
"think": {
"provider": {
"type": "open_ai",
"model": "gpt-4o-mini",
"temperature": 0.7
},
"endpoint": { // Optional if LLM provider is open_ai, anthropic, or google. Required for 3rd party LLM providers such as groq
"url": "https://api.example.com/llm", // Required if endpoint is provided
"headers": { // Optional if an endpoint is provided
"authorization": "Bearer {{token}}"
}
},
}
// ... other settings ...Passing a custom (BYO) LLM through a Cloud Provider
Section titled “Passing a custom (BYO) LLM through a Cloud Provider”Deepgram tests against major LLM providers including OpenAI, Anthropic, and Google. When bringing your own LLM, you have two options:
- Use an OpenAI-compatible LLM service or gateway. Set
provider.typetoopen_aiand point theendpoint.urlto your service. Any LLM endpoint that conforms to the OpenAI Chat Completions API format will work, including third-party LLM gateways. - Use a custom endpoint from one of the supported major LLM providers. If you have your own contract or deployment with a supported provider (such as OpenAI, Anthropic, or Google), set the
provider.typeto match that provider and supply your ownendpoint.urlandendpoint.headers.
In both cases, configure the provider.type to one of the supported provider values and set the endpoint.url and endpoint.headers fields to the correct values for your provider or gateway.
JSON
// ... other settings ...
"think": {
"provider": {
"type": "open_ai",
"model": "gpt-4",
"temperature": 0.7
},
"endpoint": { // Required for a custom LLM
"url": "https://cloud.provider.com/llm", // Required for a custom LLM
"headers": { // Optional for a custom LLM
"authorization": "Bearer {{token}}"
}
},
}
// ... other settings ...Using multiple LLM providers
Section titled “Using multiple LLM providers”The think object accepts both a single provider and an array of providers. When you supply an array, the Voice Agent uses the providers as an ordered fallback chain: it sends each LLM request to the first provider in the list and automatically falls back to the next provider if the request fails.
How fallback works
Section titled “How fallback works”- The agent sends the request to the first provider in the array.
- If that provider returns an error or times out, the agent sends a
THINK_REQUEST_FAILEDwarning over the WebSocket and retries with the next provider. - This continues through every provider in the array.
- If all providers fail, the agent sends a
FAILED_TO_THINKerror and the turn produces no LLM response.
The fallback is per-request — each new conversational turn starts again from the first provider. Provider order matters, so place your preferred provider first and your most reliable fallback last.
Example
Section titled “Example”JSON
{
"agent": {
"think": [
{
"provider": {
"type": "open_ai",
"model": "gpt-4o-mini",
"temperature": 0.7
}
},
{
"provider": {
"type": "anthropic",
"model": "claude-haiku-4-5",
"temperature": 0.7
}
}
]
}
}Listing supported models via the API
Section titled “Listing supported models via the API”The current list of providers and models is exposed by a public API endpoint. Query it whenever you need to discover which model IDs are valid for which provider, or to programmatically build a model picker.
Request
Section titled “Request”GET https://agent.deepgram.com/v1/agent/settings/think/models
curl https://agent.deepgram.com/v1/agent/settings/think/models
import requests
url = "https://agent.deepgram.com/v1/agent/settings/think/models"
response = requests.get(url)
print(response.json())
const res = await fetch(
"https://agent.deepgram.com/v1/agent/settings/think/models",
);
const data = await res.json();
console.log(data);
package main
import (
"fmt"
"net/http"
"io"
)
func main() {
url := "https://agent.deepgram.com/v1/agent/settings/think/models"
req, _ := http.NewRequest("GET", url, nil)
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
body, _ := io.ReadAll(res.Body)
fmt.Println(res)
fmt.Println(string(body))
}require 'uri'
require 'net/http'
url = URI("https://agent.deepgram.com/v1/agent/settings/think/models")
http = Net::HTTP.new(url.host, url.port)
http.use_ssl = true
request = Net::HTTP::Get.new(url)
response = http.request(request)
puts response.read_bodyimport com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;
HttpResponse<String> response = Unirest.get("https://agent.deepgram.com/v1/agent/settings/think/models")
.asString();<?php
require_once('vendor/autoload.php');
$client = new \GuzzleHttp\Client();
$response = $client->request('GET', 'https://agent.deepgram.com/v1/agent/settings/think/models');
echo $response->getBody();using RestSharp;
var client = new RestClient("https://agent.deepgram.com/v1/agent/settings/think/models");
var request = new RestRequest(Method.GET);
IRestResponse response = client.Execute(request);import Foundation
let request = NSMutableURLRequest(url: NSURL(string: "https://agent.deepgram.com/v1/agent/settings/think/models")! as URL,
cachePolicy: .useProtocolCachePolicy,
timeoutInterval: 10.0)
request.httpMethod = "GET"
let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
if (error != nil) {
print(error as Any)
} else {
let httpResponse = response as? HTTPURLResponse
print(httpResponse)
}
})
dataTask.resume()Response (200)
Section titled “Response (200)”{
"models": [
{
"id": "gpt-5",
"name": "GPT-5",
"provider": "open_ai"
}
]
}