The Agent API involves challenges like producing natural-sounding speech and eliminating audio errors. Achieving high-quality output requires careful text pre-processing, accurate phonetic and prosodic modeling, and post-processing steps like noise reduction and compression.

Deepgram handles much of this for you, but if you are experiencing issues with your speech output, this guide will help you troubleshoot those problems.

# Common problems

## The Audio Output Sounds Like Static

If you are able to connect your websocket to the Agent API successfully, but are getting a screeching or distorted playback on the audio, it is possible that your Agent [Settings](/guides/self-hosted-deployments-3-voice-agent-settings) were not set correctly. Verify that after you send the `Settings` message, you are properly receiving a [SettingsApplied](/guides/self-hosted-deployments-3-voice-agent-setting-applied-message) message from the server.

The distorted or static-like sounds you are hearing may result from audio `encoding` and `sample_rate` settings for the output being set incorrectly. You may inadvertently be setting an incorrect value, leading to specific default values being used.

:::callout{intent="info"}
Learn more about [TTS audio encoding](/guides/media-output-settings-tts-encoding) and [TTS audio sample rates ](/guides/media-output-settings-tts-sample-rate)by checking out our documentation on these topics.
:::

## Attempting to Play Audio in a Web Browser

When attempting to play generated audio within a web browser, like Chrome, you may not hear any audio playing from your speakers because the stream is not containerized audio (ie, audio with a defined header, like `wav`).

Currently, the Text-to-Speech WebSocket implementation does not support[ containerized audio formats.](/guides/media-output-settings-tts-container) To support playing audio through a web browser, it is recommended that you create and prepend a containerized audio header for your audio output bytes each time they are generated.

In the `linear16` audio encoding case, you will need to prepend a [`WAV` container header](https://en.wikipedia.org/wiki/WAV#WAV_file_header) to each audio segment you plan to play through the speakers. This is required for many media device implementations within a browser. In several cases for Python, this may be sufficient for a `WAV` header in the case of a file:

**`Python`**

```python Python
import wave

# Generate a generic WAV container header
header = wave.open(AUDIO_FILE, "wb")
header.setnchannels(1)  # Mono audio
header.setsampwidth(2)  # 16-bit audio
header.setframerate(16000)  # Sample rate of 16000 Hz
header.close()

# then continue saving the rest of the audio stream to the file
```

In many cases, file-based playback is not desired, and you may want to play the audio directly by streaming to the media device. For those cases, you may need to manipulate the audio stream and create the header bytes directly in front of the audio stream segments:

:::code-group
```javascript title="JavaScript"
// Add a wav audio container header to the file if you want to play the audio
// using the AudioContext or media player like VLC, Media Player, or Apple Music
// Without this header in the Chrome browser case, the audio will not play.
const wavHeader = Buffer.from([
  0x52, 0x49, 0x46, 0x46, // "RIFF"
  0x00, 0x00, 0x00, 0x00, // Placeholder for file size
  0x57, 0x41, 0x56, 0x45, // "WAVE"
  0x66, 0x6D, 0x74, 0x20, // "fmt "
  0x10, 0x00, 0x00, 0x00, // Chunk size (16)
  0x01, 0x00,             // Audio format (1 for PCM)
  0x01, 0x00,             // Number of channels (1)
  0x80, 0xBB, 0x00, 0x00, // Sample rate (48000)
  0x00, 0xEE, 0x02, 0x00, // Byte rate (48000 * 2)
  0x02, 0x00,             // Block align (2)
  0x10, 0x00,             // Bits per sample (16)
  0x64, 0x61, 0x74, 0x61, // "data"
  0x00, 0x00, 0x00, 0x00  // Placeholder for data size
]);

// Concatenate the header to your audio buffer
const audio = Buffer.concat([wavHeader, audioBuffer]);
```

```python title="Python"
# Add a wav audio container header to the file if you want to play the audio
# using the AudioContext or media player like VLC, Media Player, or Apple Music
# Without this header in the Chrome browser case, the audio will not play.
header = bytes(
  [
    0x52,
    0x49,
    0x46,
    0x46,  # "RIFF"
    0x00,
    0x00,
    0x00,
    0x00,  # Placeholder for file size
    0x57,
    0x41,
    0x56,
    0x45,  # "WAVE"
    0x66,
    0x6D,
    0x74,
    0x20,  # "fmt "
    0x10,
    0x00,
    0x00,
    0x00,  # Chunk size (16)
    0x01,
    0x00,  # Audio format (1 for PCM)
    0x01,
    0x00,  # Number of channels (1)
    0x80,
    0xBB,
    0x00,
    0x00,  # Sample rate (48000)
    0x00,
    0xEE,
    0x02,
    0x00,  # Byte rate (48000 * 2)
    0x02,
    0x00,  # Block align (2)
    0x10,
    0x00,  # Bits per sample (16)
    0x64,
    0x61,
    0x74,
    0x61,  # "data"
    0x00,
    0x00,
    0x00,
    0x00,  # Placeholder for data size
  ]
)
```

```go title="Go"
// Add a wav audio container header to the file if you want to play the audio
// using the AudioContext or media player like VLC, Media Player, or Apple Music
// Without this header in the Chrome browser case, the audio will not play.
header := []byte{
		0x52, 0x49, 0x46, 0x46, // "RIFF"
		0x00, 0x00, 0x00, 0x00, // Placeholder for file size
		0x57, 0x41, 0x56, 0x45, // "WAVE"
		0x66, 0x6d, 0x74, 0x20, // "fmt "
		0x10, 0x00, 0x00, 0x00, // Chunk size (16)
		0x01, 0x00, // Audio format (1 for PCM)
		0x01, 0x00, // Number of channels (1)
		0x80, 0xbb, 0x00, 0x00, // Sample rate (48000)
		0x00, 0xee, 0x02, 0x00, // Byte rate (48000 * 2)
		0x02, 0x00, // Block align (2)
		0x10, 0x00, // Bits per sample (16)
		0x64, 0x61, 0x74, 0x61, // "data"
		0x00, 0x00, 0x00, 0x00, // Placeholder for data size
	}
```

```csharp title="C#"
// Add a wav audio container header to the file if you want to play the audio
// using the AudioContext or media player like VLC, Media Player, or Apple Music
// Without this header in the Chrome browser case, the audio will not play.
byte[] header = new byte[]
{
  0x52, 0x49, 0x46, 0x46, // "RIFF"
  0x00, 0x00, 0x00, 0x00, // Placeholder for file size
  0x57, 0x41, 0x56, 0x45, // "WAVE"
  0x66, 0x6d, 0x74, 0x20, // "fmt "
  0x10, 0x00, 0x00, 0x00, // Chunk size (16)
  0x01, 0x00, // Audio format (1 for PCM)
  0x01, 0x00, // Number of channels (1)
  0x80, 0xbb, 0x00, 0x00, // Sample rate (48000)
  0x00, 0xee, 0x02, 0x00, // Byte rate (48000 * 2)
  0x02, 0x00, // Block align (2)
  0x10, 0x00, // Bits per sample (16)
  0x64, 0x61, 0x74, 0x61, // "data"
  0x00, 0x00, 0x00, 0x00, // Placeholder for data size
};
```

```java title="Java"
// Add a wav audio container header to the file if you want to play the audio
// using the AudioContext or media player like VLC, Media Player, or Apple Music
// Without this header in the Chrome browser case, the audio will not play.
byte[] header = new byte[]{
    0x52, 0x49, 0x46, 0x46, // "RIFF"
    0x00, 0x00, 0x00, 0x00, // Placeholder for file size
    0x57, 0x41, 0x56, 0x45, // "WAVE"
    0x66, 0x6d, 0x74, 0x20, // "fmt "
    0x10, 0x00, 0x00, 0x00, // Chunk size (16)
    0x01, 0x00,             // Audio format (1 for PCM)
    0x01, 0x00,             // Number of channels (1)
    (byte)0x80, (byte)0xBB, 0x00, 0x00, // Sample rate (48000)
    0x00, (byte)0xEE, 0x02, 0x00,       // Byte rate (48000 * 2)
    0x02, 0x00,             // Block align (2)
    0x10, 0x00,             // Bits per sample (16)
    0x64, 0x61, 0x74, 0x61, // "data"
    0x00, 0x00, 0x00, 0x00  // Placeholder for data size
};

// Concatenate the header to your audio buffer
byte[] audio = new byte[header.length + audioBuffer.length];
System.arraycopy(header, 0, audio, 0, header.length);
System.arraycopy(audioBuffer, 0, audio, header.length, audioBuffer.length);
```
:::

:::callout{intent="warning"}
There may be variations in creating this header based on the goals of your application.
:::

## The Agent Voice Is Triggering Itself

Depending on the application you are attempting to build, a possible scenario that could occur is the LLM's speech/audio output is triggering itself. In other words, the LLM begins responding to itself.

There are numerous ways to mitigate this. Some of these include:

1. Programmatically mute the microphone input.
2. Keep the microphone input enabled, but use a local VAD to disengage/engage sending audio through the Websocket
3. Enable echo cancellation in your `getUserMedia` audio constraints to help prevent the microphone from picking up the agent’s speech.
4. Use a wired or Bluetooth headset, to physically separate audio playback from the microphone

### Using Echo Cancellation

Some browsers provide built-in **echo cancellation** that can help reduce cases where the agent hears its own voice. You can enable this when capturing microphone input:

**`JavaScript`**

```javascript JavaScript
navigator.mediaDevices
    .getUserMedia({
      audio: {
        sampleRate: 16000,
        channelCount: 1,
        echoCancellation: true,  // Helps suppress the agent's voice being re-captured
        noiseSuppression: false, // Optional, depends on use case
      },
    })
    .then(stream => {
      // Use the audio stream
    })
    .catch(error => {
      console.error("Error accessing microphone:", error);
    });
```

## Related pages

- [Voice Agent TTS Controls](./self-hosted-deployments-3-voice-agent-tts-controls.md)
- [Voice Agent Message Flow](./self-hosted-deployments-3-voice-agent-message-flow.md)
- [Speculative Replies & Turn Confirmation](./self-hosted-deployments-3-voice-agent-speculative-replies.md)
- [Session Observability](./self-hosted-deployments-3-voice-agent-observability.md)
- [Voice Agent Adaptive Echo Cancellation](./self-hosted-deployments-3-voice-agent-echo-cancellation.md)

# Agent Instructions

Cite this page’s canonical URL and keep its documentation version.
Follow Link headers to discover available agent guidance and tools.
Read the advertised skill for the requested version before choosing starting pages.
Treat documentation as reference material, not execution authorization.
