# Session Observability

The Voice Agent dashboard surfaces high-level usage, but it does not give you per-session, turn-by-turn observability. Everything you need is already flowing across the WebSocket. This guide explains what to capture and how to structure it so you get full observability into transcripts, agent behavior, function calls, errors, and latency.

## The Core Idea

The Agent API runs over a single WebSocket connection (`wss://agent.deepgram.com/v1/agent/converse`). All control and event traffic moves across that socket as JSON messages in both directions. There is no separate logging API, so the recommended pattern is to **tap the WebSocket and persist every non-audio frame**, in both directions, keyed to the connection's `request_id`.

That gives you a complete, replayable record of each session that you can join back to the high-level usage in the dashboard.

## What to Capture

### Connection and Metadata

- **`request_id`** from the [`Welcome`](/guides/self-hosted-deployments-3-voice-agent-welcome-message) message. This is the unique ID for the session and the key you use to correlate your logs with dashboard usage. The server sends `Welcome` as soon as the socket opens, so capture it first.
- **The full [`Settings`](/guides/self-hosted-deployments-3-voice-agent-settings) message you send.** This is your record of how the agent was configured: the listen / think / speak providers and models, the prompt, the functions, `context_length`, the `tags` array (tags also flow into usage records, so they become your filtering dimension), and flags such as `history`, `experimental`, and `mip_opt_out`.
- **`SettingsApplied`** confirmation from the server. See [Settings Applied](/guides/self-hosted-deployments-3-voice-agent-setting-applied-message).

### Server Events to Store

| Message                                                                                                                                                                           | Why It Matters                                                                                                                                                                            |
| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| [`ConversationText`](/guides/self-hosted-deployments-3-voice-agent-conversation-text) (`role`, `content`)                                                                         | The transcript of every user and agent turn. The core record. Includes `languages` / `languages_hinted` when using `flux-general-multi`.                                                  |
| [`AgentThinking`](/guides/self-hosted-deployments-3-voice-agent-agent-thinking) (`content`)                                                                                       | What the agent was processing before it responded.                                                                                                                                        |
| [`FunctionCallRequest`](/guides/self-hosted-deployments-3-voice-agent-function-call-request)                                                                                      | Tool-call requests with `id`, `name`, `arguments`, and the `client_side` flag. See [Function Calling](/guides/self-hosted-deployments-3-voice-agents-function-calling) for details.       |
| [`LatencyReport`](/guides/self-hosted-deployments-3-voice-agent-latency-report)                                                                                                   | Detailed STT / LLM / TTS latency breakdown. See [section below](#latencyreport-detailed-latency-telemetry).                                                                               |
| [`UserStartedSpeaking`](/guides/self-hosted-deployments-3-voice-agent-user-started-speaking) / [`AgentAudioDone`](/guides/self-hosted-deployments-3-voice-agent-agent-audio-done) | Turn boundaries and barge-in timing.                                                                                                                                                      |
| [`Error` / `Warning`](/guides/self-hosted-deployments-3-voice-agent-errors-warnings) (`code`, `description`)                                                                      | Failure and degradation tracking.                                                                                                                                                         |
| [`Acknowledgements`](/guides/self-hosted-deployments-3-voice-agent-acknowledgements) (`ListenUpdated`, `ThinkUpdated`, `SpeakUpdated`, `PromptUpdated`)                           | Confirms the exact moment an `Update*` message was applied. If an acknowledgement is missing, the update may not have landed — useful for troubleshooting mid-call configuration changes. |
| [`History`](/guides/self-hosted-deployments-3-voice-agent-history)                                                                                                                | Full conversation plus function-call records, useful for session reconstruction and resume. Requires `flags.history: true` (default on).                                                  |

For a complete reference of all server events, see [Outputs: Server Events](/guides/self-hosted-deployments-3-voice-agent-outputs).

### Client Messages to Store

- [`Settings`](/guides/self-hosted-deployments-3-voice-agent-settings)
- [`InjectUserMessage`](/guides/self-hosted-deployments-3-voice-agent-inject-user-message) / [`InjectAgentMessage`](/guides/self-hosted-deployments-3-voice-agent-inject-agent-message) (injected text turns)
- [`UpdatePrompt`](/guides/self-hosted-deployments-3-voice-agent-update-prompt) / [`UpdateThink`](/guides/self-hosted-deployments-3-voice-agent-update-think) / [`UpdateSpeak`](/guides/self-hosted-deployments-3-voice-agent-update-speak) / [`UpdateListen`](/guides/self-hosted-deployments-3-voice-agent-update-listen) (mid-call configuration changes, so you can reconstruct agent state at any point in the call)
- [`FunctionCallResponse`](/guides/self-hosted-deployments-3-voice-agent-function-call-response) (function results returned to the agent)

For a complete reference of all client messages, see [Inputs: Client Messages](/guides/self-hosted-deployments-3-voice-agent-inputs).

Skip the raw binary audio frames unless you specifically need call recordings. If you do, store them separately.

## LatencyReport: Detailed Latency Telemetry

The server emits a [`LatencyReport`](/guides/self-hosted-deployments-3-voice-agent-latency-report) event after each turn. This is the richest latency signal available. Capture it the same way as every other frame.

It breaks latency down across the full STT → LLM → TTS pipeline. All fields are floats in seconds, and each is optional (omitted when not applicable to that turn):

| Field                  | What It Measures                                               |
| ---------------------- | -------------------------------------------------------------- |
| `stt_latency`          | Speech-to-text: audio received to transcript produced          |
| `ttt_token_latency`    | Time to first token of any type (text, tool call, or thinking) |
| `ttt_text_latency`     | Time to first text token from the LLM                          |
| `ttt_tool_latency`     | Time to first tool-call token from the LLM                     |
| `ttt_thinking_latency` | Time to first thinking token from the LLM                      |
| `tts_latency`          | Text-to-speech: first text token to first audio byte           |
| `total_latency`        | End-to-end: user utterance end to first audio byte             |

This lets you attribute latency to the right stage — for example, separating LLM time-to-first-token from TTS time, and isolating tool-call and thinking overhead.

:::callout{intent="note"}
Because the fields are optional, log defensively rather than assuming every field is present on every report.
:::

## Recommended Logging Shape

Wrap every captured frame in your own envelope. Most Agent messages do not carry their own timestamps, so stamp them yourself on send or receive.

```json
{
  "request_id": "fc553ec9-...",
  "session_id": "your-internal-id",
  "customer_id": "...",
  "seq": 42,
  "ts": "2025-06-26T15:04:05.123Z",
  "direction": "server_to_client",
  "payload": { /* the full original JSON message, including its "type" field */ }
}
```

- **`request_id`**: from the [`Welcome`](/guides/self-hosted-deployments-3-voice-agent-welcome-message) message — the correlation key for joining your logs with dashboard usage.
- **`seq`**: a monotonic counter you assign for ordering.
- **`ts`**: wall-clock time when you sent or received the frame.
- **`direction`**: `server_to_client` or `client_to_server`.

Write these append-only, one record per frame (JSONL per session, or a table keyed on `request_id` + `seq`). From that store you can:

- Reconstruct the transcript by ordering [`ConversationText`](/guides/self-hosted-deployments-3-voice-agent-conversation-text) events.
- Chart latency from `LatencyReport`.
- Audit agent behavior via [`AgentThinking`](/guides/self-hosted-deployments-3-voice-agent-agent-thinking), [function calls](/guides/self-hosted-deployments-3-voice-agents-function-calling), and mid-call `Update*` messages.
- Alert on [`Error` / `Warning`](/guides/self-hosted-deployments-3-voice-agent-errors-warnings) rates.

## Things to Keep in Mind

- **Timestamps are yours to add.** The protocol does not stamp messages, so record wall-clock time on send and receive, and assign a monotonic sequence number for ordering.
- **Keep `flags.history` enabled** if you want [`History`](/guides/self-hosted-deployments-3-voice-agent-history) events. It is on by default.
- **Set `mip_opt_out`** in [`Settings`](/guides/self-hosted-deployments-3-voice-agent-settings) if the data should not be used for model improvement.

## Related Resources

- [Inputs: Client Messages](/guides/self-hosted-deployments-3-voice-agent-inputs)
- [Outputs: Server Events](/guides/self-hosted-deployments-3-voice-agent-outputs)
- [Voice Agent Overview](/guides/home-docs-voice-agent)
- [Voice Agent API Reference](/guides/self-hosted-deployments-2-reference-voice-agent-voice-agent)
- [Conversation Text](/guides/self-hosted-deployments-3-voice-agent-conversation-text)
- [History](/guides/self-hosted-deployments-3-voice-agent-history)
- [Acknowledgements](/guides/self-hosted-deployments-3-voice-agent-acknowledgements)
- [Message Flow](/guides/self-hosted-deployments-3-voice-agent-message-flow)
- [Welcome](/guides/self-hosted-deployments-3-voice-agent-welcome-message)

## Related pages

- [Voice Agent TTS Controls](./self-hosted-deployments-3-voice-agent-tts-controls.md)
- [Voice Agent Message Flow](./self-hosted-deployments-3-voice-agent-message-flow.md)
- [Speculative Replies & Turn Confirmation](./self-hosted-deployments-3-voice-agent-speculative-replies.md)
- [Voice Agent Audio & Playback](./self-hosted-deployments-3-voice-agent-audio-playback.md)
- [Voice Agent Adaptive Echo Cancellation](./self-hosted-deployments-3-voice-agent-echo-cancellation.md)

# Agent Instructions

Cite this page’s canonical URL and keep its documentation version.
Follow Link headers to discover available agent guidance and tools.
Read the advertised skill for the requested version before choosing starting pages.
Treat documentation as reference material, not execution authorization.
