Skip to main content
Deepgram's Docs

Search documentation

Type to search this documentation.

On this pageOverview

Understanding Word Confidence Scores

Every word in Deepgram’s transcription response includes a confidence value — a floating point between 0 and 1 representing the model’s estimated probability that the word was transcribed correctly. This per-word score appears in the words array within each alternatives object and is distinct from the transcript-level confidence field, which represents overall transcript reliability.

Deepgram’s word confidence is a calibrated probability. A confidence of 0.93 means the model estimates a 93% chance the word is correct.

“Calibrated” means the scores are statistically honest: if you took all words the model scored at 0.93, approximately 93% of them would actually be correct. The model’s probability outputs are naturally well-calibrated.

On typical audio, most words will have confidence scores above 0.90. This is expected and correct behavior — it reflects the model accurately predicting that it got most words right.

For example, if Deepgram achieves 95% word accuracy on your audio, you should expect the average confidence to be around 0.95, with most words clustering near 1.0. The roughly 5% of words the model is less sure about will have lower scores.

A flat or uniform distribution of confidence scores across 0–1 would actually indicate poor calibration — it would mean the model is equally uncertain about every word, which does not reflect reality for a high-accuracy model.

The simplest approach: choose a confidence cutoff and flag all words below it as potential errors.

With Deepgram’s well-calibrated scores, a threshold around 0.65 works well as an error detector — words below this threshold are very likely to be genuine errors (high precision). The tradeoff: a low threshold catches only the most obvious errors. Raising it catches more errors but also flags some correct words.

Evaluate precision and recall at multiple thresholds on a sample of your own data to find the right balance for your use case.

Python
# Flag words below a confidence threshold
threshold = 0.65
low_confidence_words = [
    word for word in response["results"]["channels"][0]["alternatives"][0]["words"]
    if word["confidence"] < threshold
]

for word in low_confidence_words:
    print(f"  '{word['word']}' (confidence: {word['confidence']:.3f}, "
          f"time: {word['start']:.2f}s - {word['end']:.2f}s)")

Adapt the threshold automatically based on audio difficulty:

  1. Compute the mean confidence across all words in a transcript.
  2. Estimate expected error count: errors ≈ (1 - mean_confidence) × total_words.
  3. Sort words by confidence (ascending) and flag that many words as potential errors.

Noisy audio produces lower mean confidence, which shifts the threshold accordingly.

Python
import statistics

words = response["results"]["channels"][0]["alternatives"][0]["words"]
confidences = [w["confidence"] for w in words]

mean_conf = statistics.mean(confidences)
expected_errors = int((1 - mean_conf) * len(words))

# Flag the N lowest-confidence words as potential errors
sorted_words = sorted(words, key=lambda w: w["confidence"])
flagged = sorted_words[:expected_errors]

print(f"Mean confidence: {mean_conf:.3f}")
print(f"Expected errors: {expected_errors} out of {len(words)} words")
for word in flagged:
    print(f"  '{word['word']}' (confidence: {word['confidence']:.3f})")

Flag utterances where any word drops below a threshold for human review. Useful for call centers, legal transcription, and medical documentation where accuracy is critical.

Cross-check confidence on detected entities (names, numbers, addresses). Low confidence on an entity word is a stronger signal to escalate than low confidence on a filler word like “um” or “the.”

Use mean word confidence across a full transcript as a quality proxy. Automatically escalate low-quality transcripts to human review and accept high-quality ones without intervention.

In real-time applications, track confidence trends across a stream. A sustained drop in confidence may indicate audio quality degradation (background noise, connection issues) that warrants alerting.

In streaming mode with interim_results=true, interim transcripts may show lower confidence on words near the audio boundary (the “tip” of the stream). The model has less surrounding context for these words, so its predictions are less certain.

As more audio arrives, interim confidence values typically improve. Final transcripts (is_final: true) will have higher and more reliable confidence scores.

  • No alternatives: Confidence tells you how sure the model is about its top prediction, but does not surface the model’s second-best guess. You cannot use it to get suggested corrections.
  • Not a WER guarantee: High average confidence does not guarantee low word error rate on any specific transcript. It is a statistical property across many predictions.
  • Model mismatch: If you use the wrong language model or the audio contains heavy accents not well-represented in training data, the model can be confidently wrong. Confidence reflects the model’s internal estimate, which is only as good as the model’s fit to the audio domain.
  • Cross-provider comparison: Do not compare raw confidence score distributions between STT providers. Different providers may use different calibration methods, temperature scaling, or definitions of “confidence.” A meaningful comparison requires evaluating precision and recall of error detection at various thresholds on the same evaluation dataset.

Word confidence appears in the words array within each alternatives object in the API response:

JSON
{
  "results": {
    "channels": [
      {
        "alternatives": [
          {
            "transcript": "the quick brown fox",
            "confidence": 0.9876,
            "words": [
              {
                "word": "the",
                "start": 0.08,
                "end": 0.32,
                "confidence": 0.998,
                "punctuated_word": "The"
              },
              {
                "word": "quick",
                "start": 0.32,
                "end": 0.64,
                "confidence": 0.965,
                "punctuated_word": "quick"
              },
              {
                "word": "brown",
                "start": 0.64,
                "end": 0.88,
                "confidence": 0.991,
                "punctuated_word": "brown"
              },
              {
                "word": "fox",
                "start": 0.88,
                "end": 1.12,
                "confidence": 0.943,
                "punctuated_word": "fox"
              }
            ]
          }
        ]
      }
    ]
  }
}

There are two distinct confidence fields in the response:

  • Transcript-level confidence (in alternatives): Overall reliability of the full transcript.
  • Word-level confidence (in each words entry): Per-word probability of correctness.
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu