Understanding Word Confidence Scores
Every word in Deepgram’s transcription response includes a confidence value — a floating point between 0 and 1 representing the model’s estimated probability that the word was transcribed correctly. This per-word score appears in the words array within each alternatives object and is distinct from the transcript-level confidence field, which represents overall transcript reliability.
What Confidence Means
Section titled “What Confidence Means”Deepgram’s word confidence is a calibrated probability. A confidence of 0.93 means the model estimates a 93% chance the word is correct.
“Calibrated” means the scores are statistically honest: if you took all words the model scored at 0.93, approximately 93% of them would actually be correct. The model’s probability outputs are naturally well-calibrated.
Why Most Scores Are High
Section titled “Why Most Scores Are High”On typical audio, most words will have confidence scores above 0.90. This is expected and correct behavior — it reflects the model accurately predicting that it got most words right.
For example, if Deepgram achieves 95% word accuracy on your audio, you should expect the average confidence to be around 0.95, with most words clustering near 1.0. The roughly 5% of words the model is less sure about will have lower scores.
A flat or uniform distribution of confidence scores across 0–1 would actually indicate poor calibration — it would mean the model is equally uncertain about every word, which does not reflect reality for a high-accuracy model.
Using Confidence for Error Detection
Section titled “Using Confidence for Error Detection”Fixed Threshold
Section titled “Fixed Threshold”The simplest approach: choose a confidence cutoff and flag all words below it as potential errors.
With Deepgram’s well-calibrated scores, a threshold around 0.65 works well as an error detector — words below this threshold are very likely to be genuine errors (high precision). The tradeoff: a low threshold catches only the most obvious errors. Raising it catches more errors but also flags some correct words.
Evaluate precision and recall at multiple thresholds on a sample of your own data to find the right balance for your use case.
# Flag words below a confidence threshold
threshold = 0.65
low_confidence_words = [
word for word in response["results"]["channels"][0]["alternatives"][0]["words"]
if word["confidence"] < threshold
]
for word in low_confidence_words:
print(f" '{word['word']}' (confidence: {word['confidence']:.3f}, "
f"time: {word['start']:.2f}s - {word['end']:.2f}s)")Dynamic Threshold
Section titled “Dynamic Threshold”Adapt the threshold automatically based on audio difficulty:
- Compute the mean confidence across all words in a transcript.
- Estimate expected error count:
errors ≈ (1 - mean_confidence) × total_words. - Sort words by confidence (ascending) and flag that many words as potential errors.
Noisy audio produces lower mean confidence, which shifts the threshold accordingly.
import statistics
words = response["results"]["channels"][0]["alternatives"][0]["words"]
confidences = [w["confidence"] for w in words]
mean_conf = statistics.mean(confidences)
expected_errors = int((1 - mean_conf) * len(words))
# Flag the N lowest-confidence words as potential errors
sorted_words = sorted(words, key=lambda w: w["confidence"])
flagged = sorted_words[:expected_errors]
print(f"Mean confidence: {mean_conf:.3f}")
print(f"Expected errors: {expected_errors} out of {len(words)} words")
for word in flagged:
print(f" '{word['word']}' (confidence: {word['confidence']:.3f})")Common Use Cases
Section titled “Common Use Cases”QA and Compliance Review
Section titled “QA and Compliance Review”Flag utterances where any word drops below a threshold for human review. Useful for call centers, legal transcription, and medical documentation where accuracy is critical.
Entity Validation
Section titled “Entity Validation”Cross-check confidence on detected entities (names, numbers, addresses). Low confidence on an entity word is a stronger signal to escalate than low confidence on a filler word like “um” or “the.”
Transcript Quality Scoring
Section titled “Transcript Quality Scoring”Use mean word confidence across a full transcript as a quality proxy. Automatically escalate low-quality transcripts to human review and accept high-quality ones without intervention.
Streaming Confidence Monitoring
Section titled “Streaming Confidence Monitoring”In real-time applications, track confidence trends across a stream. A sustained drop in confidence may indicate audio quality degradation (background noise, connection issues) that warrants alerting.
Confidence in Streaming vs. Pre-recorded
Section titled “Confidence in Streaming vs. Pre-recorded”In streaming mode with interim_results=true, interim transcripts may show lower confidence on words near the audio boundary (the “tip” of the stream). The model has less surrounding context for these words, so its predictions are less certain.
As more audio arrives, interim confidence values typically improve. Final transcripts (is_final: true) will have higher and more reliable confidence scores.
Limitations
Section titled “Limitations”- No alternatives: Confidence tells you how sure the model is about its top prediction, but does not surface the model’s second-best guess. You cannot use it to get suggested corrections.
- Not a WER guarantee: High average confidence does not guarantee low word error rate on any specific transcript. It is a statistical property across many predictions.
- Model mismatch: If you use the wrong language model or the audio contains heavy accents not well-represented in training data, the model can be confidently wrong. Confidence reflects the model’s internal estimate, which is only as good as the model’s fit to the audio domain.
- Cross-provider comparison: Do not compare raw confidence score distributions between STT providers. Different providers may use different calibration methods, temperature scaling, or definitions of “confidence.” A meaningful comparison requires evaluating precision and recall of error detection at various thresholds on the same evaluation dataset.
API Reference
Section titled “API Reference”Word confidence appears in the words array within each alternatives object in the API response:
{
"results": {
"channels": [
{
"alternatives": [
{
"transcript": "the quick brown fox",
"confidence": 0.9876,
"words": [
{
"word": "the",
"start": 0.08,
"end": 0.32,
"confidence": 0.998,
"punctuated_word": "The"
},
{
"word": "quick",
"start": 0.32,
"end": 0.64,
"confidence": 0.965,
"punctuated_word": "quick"
},
{
"word": "brown",
"start": 0.64,
"end": 0.88,
"confidence": 0.991,
"punctuated_word": "brown"
},
{
"word": "fox",
"start": 0.88,
"end": 1.12,
"confidence": 0.943,
"punctuated_word": "fox"
}
]
}
]
}
]
}
}There are two distinct confidence fields in the response:
- Transcript-level
confidence(inalternatives): Overall reliability of the full transcript. - Word-level
confidence(in eachwordsentry): Per-word probability of correctness.
Related Resources
Section titled “Related Resources”- Interim Results — How interim and final transcripts work in streaming.
- Utterances — Segment speech into meaningful semantic units.
- Pre-recorded Audio Getting Started — Transcribe audio files with Deepgram.
- Streaming Audio Getting Started — Real-time transcription with WebSockets.