Appearance
For clean Markdown of any page, append .md to the page URL. For a complete documentation index, see For full documentation content, see For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at
Speaker Identification
Identify speakers by name or role in your transcript
For the complete documentation index, see llms.txt
Feature: Speaker Identification - replace generic "Speaker A/B" labels with real names or roles.
Supported models: Universal-3 Pro (universal-3-pro), Universal-2 (universal-2)
Supported regions: US and EU
Supported languages: 90+ languages. See Supported Languages for the full list.
Prerequisite: Speaker Diarization must be enabled (speaker_labels: true).
Key API parameters (nested under speech_understanding.request.speaker_identification):
speaker_type(string, required) -"name"or"role"known_values(array) - List of speaker names or roles (each max 35 chars). Use this ORspeakers, not both.speakers(array) - Speaker objects with metadata. Each hasnameorrole(depending onspeaker_type), optionaldescription, and any custom properties.
Common role combinations:
["Agent", "Customer"]- Customer service calls["Interviewer", "Interviewee"]- Interview recordings["Host", "Guest"]- Podcast or show recordings
cURL quickstart (identify by name):
curl -X POST "" \
-H "Authorization: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio_url": "YOUR_AUDIO_URL",
"speech_models": ["universal-3-pro", "universal-2"],
"speaker_labels": true,
"speech_understanding": {
"request": {
"speaker_identification": {
"speaker_type": "name",
"known_values": ["Michel Martin", "Peter DeCarlo"]
}
}
}
}'Poll GET /v2/transcript/{id} until status is completed. The utterances array will have identified speaker names/roles instead of generic labels. The response includes speech_understanding.response.speaker_identification.mapping showing the label-to-name mapping.
US & EU
Overview
Replace generic "Speaker A" and "Speaker B" labels with real names or roles, no voice enrollment needed. Speaker Identification uses conversation content to infer who's speaking and applies the identifiers you provide.
Example transformation:
Before:
txt
Speaker A: Good morning, and welcome to the show.
Speaker B: Thanks for having me.
Speaker A: Let's dive into today's topic...After (by name):
txt
Michel Martin: Good morning, and welcome to the show.
Peter DeCarlo: Thanks for having me.
Michel Martin: Let's dive into today's topic...After (by role):
txt
Interviewer: Good morning, and welcome to the show.
Interviewee: Thanks for having me.
Interviewer: Let's dive into today's topic...Speaker Identification requires Speaker Diarization. You must set speaker_labels: true in your transcription request.
To reliably identify speakers, your audio should contain clear, distinguishable voices and sufficient spoken audio from each speaker. The accuracy of Speaker Diarization depends on the quality of the audio and the distinctiveness of each speaker's voice, which will have a downstream effect on the quality of Speaker Identification.
Choosing how to identify speakers
You can identify speakers by name or by role:
- Know the speakers' names? Use
speaker_type: "name"with the names inknown_valuesorspeakers. Click here to learn more. - Know their roles but not names? Use
speaker_type: "role"with roles like"Interviewer"or"Agent"inknown_valuesorspeakers. Click here to learn more. - Need better accuracy? Use
speakerswithdescriptionfields that provide context about what each speaker typically discusses. Click here to learn more.
How to use Speaker Identification
Include the speech_understanding parameter in your transcription request to identify speakers.
Already have a completed transcript? You can add Speaker Identification to an existing transcript in a separate request.
Identify by name
To identify speakers by name, use speaker_type: "name" with a list of speaker names in known_values. This is the most common approach when you know who is speaking in the audio.
python
import requests
import time
base_url = ""
headers = {
"authorization": "<YOUR_API_KEY>"
}
# Need to transcribe a local file? Learn more here:
upload_url = ""
# Configure transcript with speaker identification
data = {
"audio_url": upload_url,
"speech_models": ["universal-3-pro", "universal-2"],
"language_detection": True,
"speaker_labels": True,
"speech_understanding": {
"request": {
"speaker_identification": {
"speaker_type": "name",
"known_values": ["Michel Martin", "Peter DeCarlo"] # Change these values to match the names of the speakers in your file
}
}
}
}
# Submit the transcription request
response = requests.post(base_url + "/v2/transcript", headers=headers, json=data)
transcript_id = response.json()["id"]
polling_endpoint = base_url + f"/v2/transcript/{transcript_id}"
# Poll for transcription results
while True:
transcript = requests.get(polling_endpoint, headers=headers).json()
if transcript["status"] == "completed":
break
elif transcript["status"] == "error":
raise RuntimeError(f"Transcription failed: {transcript['error']}")
else:
time.sleep(3)
# Access the results and print utterances to the terminal
for utterance in transcript["utterances"]:
print(f"{utterance['speaker']}: {utterance['text']}"){/*
python
import assemblyai as aai
aai.settings.api_key = "<YOUR_API_KEY>"
# Need to transcribe a local file? Learn more here:
audio_url = ""
# Configure transcript with speaker identification
config = aai.TranscriptionConfig(
speaker_labels=True,
speech_understanding=aai.SpeechUnderstandingConfig(
speaker_identification=aai.SpeakerIdentificationConfig(
speaker_type="name",
known_values=["Michel Martin", "Peter DeCarlo"] # Change these values to match the names of the speakers in your file
)
)
)
transcriber = aai.Transcriber()
transcript = transcriber.transcribe(audio_url, config)
# Access the results and print utterances to the terminal
for utterance in transcript.utterances:
print(f"{utterance.speaker}: {utterance.text}")*/}
javascript
const baseUrl = "";
const headers = {
"authorization": "<YOUR_API_KEY>",
"content-type": "application/json"
};
// Need to transcribe a local file? Learn more here:
const uploadUrl = "";
// Configure transcript with speaker identification
const data = {
audio_url: uploadUrl,
speech_models: ["universal-3-pro", "universal-2"],
language_detection: true,
speaker_labels: true,
speech_understanding: {
request: {
speaker_identification: {
speaker_type: "name",
known_values: ["Michel Martin", "Peter DeCarlo"] // Change these values to match the names of the speakers in your file
}
}
}
};
async function main() {
// Submit the transcription request
const response = await fetch(`${baseUrl}/v2/transcript`, {
method: "POST",
headers: headers,
body: JSON.stringify(data)
});
const { id: transcriptId } = await response.json();
const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`;
// Poll for transcription results
while (true) {
const pollingResponse = await fetch(pollingEndpoint, { headers });
const transcript = await pollingResponse.json();
if (transcript.status === "completed") {
// Access the results and print utterances to the console
for (const utterance of transcript.utterances) {
console.log(`${utterance.speaker}: ${utterance.text}`);
}
break;
} else if (transcript.status === "error") {
throw new Error(`Transcription failed: ${transcript.error}`);
} else {
await new Promise(resolve => setTimeout(resolve, 3000));
}
}
}
main().catch(console.error);python
import assemblyai as aai
aai.settings.api_key = "<YOUR_API_KEY>"
# Need to transcribe a local file? Learn more here:
audio_url = ""
# Configure transcript with speaker identification
config = aai.TranscriptionConfig(
speech_models=["universal-3-pro", "universal-2"],
language_detection=True,
speaker_labels=True,
speech_understanding=aai.SpeechUnderstandingRequest(
request=aai.SpeechUnderstandingFeatureRequests(
speaker_identification=aai.SpeakerIdentificationRequest(
speaker_type="name",
known_values=["Michel Martin", "Peter DeCarlo"] # Change these values to match the names of the speakers in your file
)
)
)
)
transcriber = aai.Transcriber()
transcript = transcriber.transcribe(audio_url, config)
# Access the results and print utterances to the terminal
for utterance in transcript.utterances:
print(f"{utterance.speaker}: {utterance.text}")javascript
import { AssemblyAI } from "assemblyai";
const client = new AssemblyAI({
apiKey: "<YOUR_API_KEY>"
});
// Need to transcribe a local file? Learn more here:
const audioUrl = "";
// Configure transcript with speaker identification
const params = {
audio: audioUrl,
speech_models: ["universal-3-pro", "universal-2"],
language_detection: true,
speaker_labels: true,
speech_understanding: {
request: {
speaker_identification: {
speaker_type: "name",
known_values: ["Michel Martin", "Peter DeCarlo"] // Change these values to match the names of the speakers in your file
}
}
}
};
const transcript = await client.transcripts.transcribe(params);
// Access the results and print utterances to the console
for (const utterance of transcript.utterances) {
console.log(`${utterance.speaker}: ${utterance.text}`);
}{/*
javascript
import { AssemblyAI } from "assemblyai";
const client = new AssemblyAI({
apiKey: "<YOUR_API_KEY>"
});
// Need to transcribe a local file? Learn more here:
const audioUrl = "";
// Configure transcript with speaker identification
const config = {
audio_url: audioUrl,
speaker_labels: true,
speech_understanding: {
speaker_identification: {
speaker_type: "name",
known_values: ["Michel Martin", "Peter DeCarlo"] // Change these values to match the names of the speakers in your file
}
}
};
const transcript = await client.transcripts.transcribe(config);
// Access the results and print utterances to the console
for (const utterance of transcript.utterances) {
console.log(`${utterance.speaker}: ${utterance.text}`);
}*/}
Identify by role
To identify speakers by role instead of name, use speaker_type: "role" with role labels in known_values. This is useful for customer service calls, interviews, or any scenario where you know the roles but not the names.
python
import requests
import time
base_url = ""
headers = {
"authorization": "<YOUR_API_KEY>"
}
# Need to transcribe a local file? Learn more here:
upload_url = ""
# Configure transcript with role-based speaker identification
data = {
"audio_url": upload_url,
"speech_models": ["universal-3-pro", "universal-2"],
"language_detection": True,
"speaker_labels": True,
"speech_understanding": {
"request": {
"speaker_identification": {
"speaker_type": "role",
"known_values": ["Interviewer", "Interviewee"] # Change these values to match the roles of the speakers in your file
}
}
}
}
# Submit the transcription request
response = requests.post(base_url + "/v2/transcript", headers=headers, json=data)
transcript_id = response.json()["id"]
polling_endpoint = base_url + f"/v2/transcript/{transcript_id}"
# Poll for transcription results
while True:
transcript = requests.get(polling_endpoint, headers=headers).json()
if transcript["status"] == "completed":
break
elif transcript["status"] == "error":
raise RuntimeError(f"Transcription failed: {transcript['error']}")
else:
time.sleep(3)
# Access the results and print utterances to the terminal
for utterance in transcript["utterances"]:
print(f"{utterance['speaker']}: {utterance['text']}")javascript
const baseUrl = "";
const headers = {
"authorization": "<YOUR_API_KEY>",
"content-type": "application/json"
};
// Need to transcribe a local file? Learn more here:
const uploadUrl = "";
// Configure transcript with role-based speaker identification
const data = {
audio_url: uploadUrl,
speech_models: ["universal-3-pro", "universal-2"],
language_detection: true,
speaker_labels: true,
speech_understanding: {
request: {
speaker_identification: {
speaker_type: "role",
known_values: ["Interviewer", "Interviewee"] // Change these values to match the roles of the speakers in your file
}
}
}
};
async function main() {
// Submit the transcription request
const response = await fetch(`${baseUrl}/v2/transcript`, {
method: "POST",
headers: headers,
body: JSON.stringify(data)
});
const { id: transcriptId } = await response.json();
const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`;
// Poll for transcription results
while (true) {
const pollingResponse = await fetch(pollingEndpoint, { headers });
const transcript = await pollingResponse.json();
if (transcript.status === "completed") {
// Access the results and print utterances to the console
for (const utterance of transcript.utterances) {
console.log(`${utterance.speaker}: ${utterance.text}`);
}
break;
} else if (transcript.status === "error") {
throw new Error(`Transcription failed: ${transcript.error}`);
} else {
await new Promise(resolve => setTimeout(resolve, 3000));
}
}
}
main().catch(console.error);python
import assemblyai as aai
aai.settings.api_key = "<YOUR_API_KEY>"
audio_url = ""
config = aai.TranscriptionConfig(
speech_models=["universal-3-pro", "universal-2"],
language_detection=True,
speaker_labels=True,
speech_understanding=aai.SpeechUnderstandingRequest(
request=aai.SpeechUnderstandingFeatureRequests(
speaker_identification=aai.SpeakerIdentificationRequest(
speaker_type="role",
known_values=["Interviewer", "Interviewee"] # Change these values to match the roles of the speakers in your file
)
)
)
)
transcriber = aai.Transcriber()
transcript = transcriber.transcribe(audio_url, config)
for utterance in transcript.utterances:
print(f"{utterance.speaker}: {utterance.text}")javascript
import { AssemblyAI } from "assemblyai";
const client = new AssemblyAI({
apiKey: "<YOUR_API_KEY>"
});
const audioUrl = "";
const params = {
audio: audioUrl,
speech_models: ["universal-3-pro", "universal-2"],
language_detection: true,
speaker_labels: true,
speech_understanding: {
request: {
speaker_identification: {
speaker_type: "role",
known_values: ["Interviewer", "Interviewee"] // Change these values to match the roles of the speakers in your file
}
}
}
};
const transcript = await client.transcripts.transcribe(params);
for (const utterance of transcript.utterances) {
console.log(`${utterance.speaker}: ${utterance.text}`);
}Common role combinations
["Agent", "Customer"]- Customer service calls["AI Assistant", "User"]- AI chatbot interactions["Support", "Customer"]- Technical support calls["Interviewer", "Interviewee"]- Interview recordings["Host", "Guest"]- Podcast or show recordings["Moderator", "Panelist"]- Panel discussions
Adding speaker metadata
For more accurate speaker identification, you can use the speakers parameter instead of known_values. The speakers parameter lets you provide additional metadata about each speaker to help the model identify speakers based on conversational context.
This is particularly useful when:
- Speakers have similar voices but distinct roles or topics
- You want to provide contextual clues about what each speaker typically discusses
- You need more precise identification in complex multi-speaker scenarios
Each speaker object must include either a name or role (depending on speaker_type). Beyond that, you can add any additional properties you want. The name and role fields are reserved as strings, but all other properties are flexible and can be any structure.
Examples in this section are shown in Python for brevity. The same speaker_identification configuration works in any language.
At its simplest, you can provide a description alongside each speaker's name or role:
python
data = {
"audio_url": upload_url,
"speaker_labels": True,
"speech_understanding": {
"request": {
"speaker_identification": {
"speaker_type": "role",
"speakers": [
{
"role": "interviewer",
"description": "Hosts the program and interviews the guests"
},
{
"role": "guest",
"description": "Answers questions from the interview"
}
]
}
}
}
}For even more fine-tuned identification, you can include any additional custom properties on each speaker object, such as company, title, department, or any other fields that help describe the speaker:
python
data = {
"audio_url": upload_url,
"speaker_labels": True,
"speech_understanding": {
"request": {
"speaker_identification": {
"speaker_type": "name",
"speakers": [
{
"name": "Michel Martin",
"description": "Hosts the program and interviews the guests",
"company": "NPR",
"title": "Host Morning Edition"
},
{
"name": "Peter DeCarlo",
"description": "Answers questions from the interview",
"company": "Johns Hopkins University",
"title": "Professor and Vice Chair of Environmental Health and Engineering"
}
]
}
}
}
}You can use the same custom properties with role-based identification by replacing name with role in each speaker object.
API reference
Request
Include the speech_understanding parameter directly in your transcription request (shown here with name-based identification):
bash
curl -X POST \
"" \
-H "Authorization: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio_url": "",
"speaker_labels": true,
"speech_understanding": {
"request": {
"speaker_identification": {
"speaker_type": "name",
"known_values": ["Michel Martin", "Peter DeCarlo"]
}
}
}
}'Request parameters
The following parameters are nested under speech_understanding.request.speaker_identification:
| Key | Type | Required? | Description |
|---|---|---|---|
speaker_type | string | Yes | The type of speakers being identified, values accepted are "name" for actual names or "role" for roles/titles. |
known_values | array | Conditional | List of speaker names or roles. Required when speaker_type is set to "role" and speakers is not provided. Optional when speaker_type is set to "name". Each value must be 35 characters or less. Use known_values or speakers, not both. |
speakers | array | Conditional | An array of speaker objects with metadata. Use as an alternative to known_values when you want to provide additional context about each speaker. You can include any additional custom properties beyond name/role and description. Use speakers or known_values, not both. |
speakers[].role | string | Conditional | The role of the speaker. Required when speaker_type is "role". |
speakers[].name | string | Conditional | The name of the speaker. Required when speaker_type is "name". |
speakers[].description | string | No | A description of the speaker to help the model identify them based on conversational context. |
speakers[]. | any | No | Any additional custom properties (e.g., company, title, department) to provide more context about the speaker. The name and role fields are reserved as strings, but all other properties are flexible. |
Response
The Speaker Identification API returns a modified version of your transcript with updated speaker labels in the utterances key.
json
{
"speech_understanding": {
"request": {
"speaker_identification": {
"speaker_type": "name",
"known_values": ["Michel Martin", "Peter DeCarlo"]
}
},
"response": {
"speaker_identification": {
"mapping": {
"A": "Michel Martin",
"B": "Peter DeCarlo"
},
"status": "success"
}
}
},
"utterances": [
{
"speaker": "Michel Martin",
"text": "Smoke from hundreds of wildfires in Canada is triggering air quality alerts...",
"start": 240,
"end": 26560,
"confidence": 0.9815734,
"words": [
{
"text": "Smoke",
"start": 240,
"end": 640,
"confidence": 0.90152997,
"speaker": "Michel Martin"
}
// ... more words
]
}
// ... more utterances
]
}Response fields
| Key | Type | Description |
|---|---|---|
speech_understanding.response.speaker_identification.mapping | object | A mapping of the original generic speaker labels (e.g., "A", "B") to the identified speaker names or roles. |
speech_understanding.response.speaker_identification.status | string | The status of the speaker identification request (e.g., "success"). |
utterances | array | A turn-by-turn temporal sequence of the transcript, where the i-th element is an object containing information about the i-th utterance in the audio file. |
utterances[i].confidence | number | The confidence score for the transcript of this utterance. |
utterances[i].end | number | The ending time, in milliseconds, of the utterance in the audio file. |
utterances[i].speaker | string | The identified speaker name or role for this utterance. |
utterances[i].start | number | The starting time, in milliseconds, of the utterance in the audio file. |
utterances[i].text | string | The transcript for this utterance. |
utterances[i].words | array | A sequential array for the words in the transcript, where the j-th element is an object containing information about the j-th word in the utterance. |
utterances[i].words[j].text | string | The text of the j-th word in the i-th utterance. |
utterances[i].words[j].start | number | The starting time for when the j-th word is spoken in the i-th utterance, in milliseconds. |
utterances[i].words[j].end | number | The ending time for when the j-th word is spoken in the i-th utterance, in milliseconds. |
utterances[i].words[j].confidence | number | The confidence score for the transcript of the j-th word in the i-th utterance. |
utterances[i].words[j].speaker | string | The identified speaker name or role who uttered the j-th word in the i-th utterance. |
With Speaker Identification, the speaker field in utterances and words contains the identified name or role (e.g., "Michel Martin" or "Agent") instead of generic labels like "A", "B", "C". All other fields (text, start, end, confidence, words) remain unchanged from the standard transcription response.