Appearance
For clean Markdown of any page, append .md to the page URL. For a complete documentation index, see For full documentation content, see For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at
Events reference
For the complete documentation index, see llms.txt
Every message exchanged over the Voice Agent API WebSocket, grouped by direction. You'll send session.update to configure, input.audio to stream mic audio, and tool.result to respond to tool calls. The server streams everything else back. For how these fit together in a typical session, see the Overview event flow.
Client → Server
input.audio
Stream PCM16 audio to the agent.
json
{
"type": "input.audio",
"audio": "<base64-encoded PCM16>"
}| Field | Type | Description |
|---|---|---|
audio | string | Base64-encoded PCM16 mono 24kHz audio |
See Audio format for the full format specification.
session.update
Configure the session. Send immediately on WebSocket connect (before session.ready). Can also be sent mid-conversation to update most fields. See Mutability after session.ready for which fields can change once the session is established.
json
{
"type": "session.update",
"session": {
"system_prompt": "You are a concise assistant.",
"greeting": "Hi! How can I help?",
"input": {
"format": { "encoding": "audio/pcm" },
"turn_detection": { "vad_threshold": 0.5 }
},
"output": {
"voice": "ivy",
"format": { "encoding": "audio/pcm" }
},
"tools": [
{
"type": "function",
"name": "get_weather",
"description": "Get weather for a city",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}
]
}
}All fields are optional. Include only what you want to set or change. After session.ready, only a subset of fields can be changed; changing greeting or session.output raises immutable_field.
| Field | Type | Description |
|---|---|---|
session.system_prompt | string | Sets the agent's personality and context |
session.greeting | string | Spoken aloud at the start of the conversation |
session.input.format | object | Input audio format (encoding). See Audio format |
session.input.keyterms | array | List of strings to boost in transcription. See Key terms |
session.input.turn_detection | object | Turn detection configuration. See Session configuration |
session.output.voice | string | The voice used for the agent's speech. See Voices |
session.output.format | object | Output audio format (encoding). See Audio format |
session.tools | array | Tool definitions. See Tool calling |
session.resume
Reconnect to an existing session using the session_id from a previous session.ready. Preserves conversation context across dropped connections.
json
{
"type": "session.resume",
"session_id": "sess_abc123"
}Sessions are preserved for 30 seconds after every disconnection before expiring. If the session has expired, the server returns a session.error with code session_not_found or session_forbidden. Start a fresh connection without session.resume.
Example. Capture session_id from session.ready on the first connection, then send session.resume as the first message when reconnecting:
python
import json
import websockets
session_id: str | None = None
async def connect():
global session_id
async with websockets.connect(URL, additional_headers={"Authorization": f"Bearer {API_KEY}"}) as ws:
# If we already have a session_id from a previous connection, resume it.
if session_id:
await ws.send(json.dumps({"type": "session.resume", "session_id": session_id}))
else:
await ws.send(json.dumps({"type": "session.update", "session": {...}}))
async for raw in ws:
event = json.loads(raw)
if event["type"] == "session.ready":
session_id = event["session_id"] # save for next reconnect
elif event["type"] == "session.error" and event["code"] in ("session_not_found", "session_forbidden"):
session_id = None # session expired - start fresh next time
# ... handle other events
# On disconnect, call connect() again within 30 seconds to resume.tool.result
Send a tool result back to the agent. Send this in the reply.done handler (not immediately in tool.call). See Tool calling.
json
{
"type": "tool.result",
"call_id": "call_abc123",
"result": "{\"temp_c\": 22, \"description\": \"Sunny\"}"
}| Field | Type | Description |
|---|---|---|
call_id | string | The call_id from the tool.call event |
result | string | JSON string containing the tool result |
Server → Client
session.ready
Session is established and ready to receive audio. Save session_id for reconnection. Start sending input.audio only after this event.
json
{
"type": "session.ready",
"session_id": "sess_abc123"
}| Field | Type | Description |
|---|---|---|
session_id | string | Always present. Save this value to reconnect with session.resume. |
session.updated
Sent after session.update is applied successfully.
json
{ "type": "session.updated" }input.speech.started
Turn detection determined the user has started speaking.
json
{ "type": "input.speech.started" }input.speech.stopped
Turn detection determined the user has stopped speaking.
json
{ "type": "input.speech.stopped" }transcript.user.delta
Partial transcript of what the user is saying, updating in real-time.
json
{
"type": "transcript.user.delta",
"text": "What's the weather in"
}transcript.user
Final transcript of the user's utterance.
json
{
"type": "transcript.user",
"text": "What's the weather in Tokyo?",
"item_id": "item_abc123"
}reply.started
Agent has begun generating a response.
json
{
"type": "reply.started",
"reply_id": "reply_abc123"
}reply.audio
A chunk of the agent's spoken response as base64 PCM16. Decode and play immediately.
json
{
"type": "reply.audio",
"data": "<base64-encoded PCM16>"
}See Audio format for playback guidance.
transcript.agent
Full text of the agent's response, sent after all audio for the response has been delivered. If the agent was interrupted, interrupted is true and text contains only what was actually spoken before the interruption.
json
{
"type": "transcript.agent",
"text": "It's currently 22°C and sunny in Tokyo.",
"reply_id": "reply_abc123",
"item_id": "item_abc123",
"interrupted": false
}| Field | Type | Description |
|---|---|---|
text | string | What the agent said (trimmed to interruption point if interrupted) |
reply_id | string | ID of the reply |
item_id | string | Conversation item ID |
interrupted | boolean | true if the user interrupted mid-response |
reply.done
Agent has finished speaking. The optional status field indicates why the reply ended.
json
{ "type": "reply.done" }json
{ "type": "reply.done", "status": "interrupted" }| Field | Type | Description |
|---|---|---|
status | string | "interrupted" if the user barged in, absent for normal completion |
tool.call
Agent wants to call a registered tool. arguments is a dict, ready to use directly as-is.
json
{
"type": "tool.call",
"call_id": "call_abc123",
"name": "get_weather",
"arguments": { "location": "Tokyo" }
}| Field | Type | Description |
|---|---|---|
call_id | string | Include this in tool.result |
name | string | Tool name to call |
arguments | object | Arguments as a dict (use directly) |
See Tool calling for the full pattern.
session.error
Session or protocol error. The payload always includes type, timestamp, code, and message. Some errors (like session.update validation failures) also include a param field naming the offending field.
json
{
"type": "session.error",
"code": "invalid_format",
"message": "Invalid message format",
"timestamp": "2025-01-01T00:00:00Z"
}Connection and handshake errors
Sent before or instead of session.ready. The WebSocket closes after these with the indicated close code.
| Code | Close code | Description |
|---|---|---|
UNAUTHORIZED | 1008 | Missing or invalid Authorization token |
FORBIDDEN | 1008 | Valid token, but insufficient permissions |
server_error | 1008 | Service at capacity (try again later) |
INTERNAL_ERROR | 1011 | Unexpected exception during connection setup |
Session resume errors
Sent when session.resume fails. The WebSocket closes after these.
| Code | Close code | Description |
|---|---|---|
session_not_found | 1008 | The session_id is unknown or the 30-second grace window expired |
session_forbidden | 1008 | The session_id belongs to a different account |
session_expired | 1008 | Session TTL elapsed during the grace window |
Agent startup errors
Sent after the WebSocket is accepted but before session.ready.
| Code | Description |
|---|---|
agent_init_failed | Voice agent worker reported initialization failure |
agent_timeout | Agent did not signal ready within 10 seconds |
Client message errors
Sent on the open socket when an inbound message is invalid. The session stays alive (except session_expired).
| Code | Description |
|---|---|
invalid_format | Bad JSON, missing or unknown type, validation failure, or missing audio field on input.audio |
invalid_audio | input.audio payload failed base64 decode or PCM conversion |
invalid_value | session.update with an invalid voice or field type |
immutable_field | session.update tried to change greeting or output after the first update was applied |
invalid_config | session.update raised a validation error |
server_error | Unexpected exception while applying session.update |
Live session errors
| Code | Close code | Description |
|---|---|---|
session_expired | 1008 | Session duration TTL reached. There is no separate "closing soon" warning event before this, so run a client-side timer if you need to wrap up gracefully. |
If the server cancels the session due to an internal error, the WebSocket closes with code 1011 without any session.error payload. In browsers, pre-handshake failures (like UNAUTHORIZED) surface as a close event with code 1006. You won't receive a session.error. Always fetch a fresh token immediately before each connection attempt.
Interruptions
When the user speaks mid-response (barge-in), the server stops the agent and emits reply.done with status: "interrupted" and transcript.agent with interrupted: true. The decision is semantic. Back-channels like "uh-huh" don't trigger an interruption. See Turn detection and interruptions for how the model decides, and Handling interruptions for the client-side flush pattern.