Skip to content

For clean Markdown of any page, append .md to the page URL. For a complete documentation index, see For full documentation content, see For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at

Speech Understanding

Extract structured insights from audio with Speech Understanding models.

For the complete documentation index, see llms.txt

Speech Understanding models analyze your transcripts to extract meaningful information like speaker identities, sentiment, topics, and summaries.

Identify and label speakers in your audio to attribute speech to the correct person.

Translate transcripts into other languages.

Customize how your transcript text is formatted.

Detect and classify entities like names, locations, and organizations in your transcripts.

Detect the sentiment of each sentence in your transcript.

Automatically segment your transcript into chapters with summaries.

Extract the most important phrases and words from your transcript.

Detect topics discussed in your audio using the IAB taxonomy.

Generate summaries of your transcripts in different formats.