Pricing reviewed September 7, 2026Review dates and sources appear on each record.

Voice & audio · September 2026

Give your project the right voice.

Narration, live conversation, transcription and music are different jobs. Match the model to the output you need, then listen to it on a real script.

Narration

Text becomes a finished voiceover. Delivery, pronunciation and consistency matter most.

Live speech

Speech arrives while someone interacts. Measure pauses, interruptions and time to audible response.

Transcription

Recorded or live speech becomes text, timestamps and speaker labels.

Music

A description or lyrics become a song or soundtrack. Arrange and mix separately from narration.

Text to speech

Six speech choices worth testing.

Choose two candidates for your specific language, voice and delivery. Provider latency claims do not measure the complete experience in your app.

Hosted · Speech

MiniMax Speech 2.8 HD

Try for polished narration and multilingual voiceovers.

Where to use it
MiniMax API

Text-to-speech with streaming, pronunciation controls and voice selection. Compare HD and Turbo on the same script.

Model identifierspeech-2.8-hd
Hosted · Speech

MiniMax Speech 2.8 Turbo

Try for interactive or repeated speech generation.

Where to use it
MiniMax API

A separate speech tier. Measure response time, pronunciation and delivered-audio cost on your workload.

Model identifierspeech-2.8-turbo
Hosted · Speech

Cartesia Sonic 3.6

Test conversational speech and responsive spoken interfaces.

Where to use it
Cartesia API / Playground

Generally available August 27, 2026; 44 languages. Pin sonic-3.6-2026-08-27 for a reproducible model snapshot.

Model identifiersonic-3.6
Hosted · Speech

Eleven v3

Expressive narration and multi-speaker dialogue.

Where to use it
ElevenLabs API / creative tools

Choose for performance and delivery; evaluate the actual voice and script.

Model identifiereleven_v3
Hosted · Speech

Eleven v3 Conversational

Expressive realtime speech synthesis.

Where to use it
ElevenLabs API

Supports realtime dialogue through a Text to Dialogue WebSocket.

Model identifiereleven_v3_conversational
Hosted · Speech

Eleven Flash v2.5

A speed-focused speech option for interactive applications.

Where to use it
ElevenLabs API

Test dates, currencies and phone numbers: text normalization differs from other models.

Model identifiereleven_flash_v2_5

Speech to text

Turn the recording into usable text.

Test names, accents, overlapping speakers and background noise before relying on an automatic transcript.

Hosted · Transcription

Scribe v2

Transcribe recordings with speaker labels and word timestamps.

Where to use it
ElevenLabs API

Speech-to-text, rather than a voice generator.

Model identifierscribe_v2
Hosted · Transcription

Scribe v2 Realtime

Streaming transcription for live audio.

Where to use it
ElevenLabs API

Evaluate transcript revisions, silence handling and noisy speech.

Model identifierscribe_v2_realtime

Music generation

Build the soundtrack separately.

Keep dialogue, music and effects on separate tracks so a change to one does not force you to regenerate everything.

Downloadable · Music

MiniMax Music 3

Turn a musical brief and lyrics into a complete song.

Where to use it
Downloadable weights / ComfyUI

Up to five minutes, 32 kHz stereo output. Review the model’s license and official workflow requirements before production use.

Model identifierMiniMaxAI/MiniMax-Music3

Choose a first voice test

For a finished voiceover, compare MiniMax Speech 2.8 HD with Eleven v3. For an interactive interface, compare Sonic 3.6, Eleven v3 Conversational, Flash v2.5 or Speech 2.8 Turbo. These are suggested experiments, not a measured ranking.

Use a short script containing a person’s name, a company name, a date, a price, an acronym and a sentence that changes emotion. Include your actual target language. First record how each should sound, then listen without looking at which model produced the result.

  1. Check pronunciation. Review names, abbreviations, dates and amounts. Save corrections with the script.
  2. Check performance. Listen for inappropriate emphasis, pauses, emotional delivery and consistency across a longer passage.
  3. Measure the whole interaction. For live use, include recognition, reasoning, network delays, first audible speech and interruption handling.
  4. Track usable output cost. Count retries, speech minutes, credits or characters, voice-plan fees and editing time. Compare the same final recording length.

A speech model is one part of a voice assistant

A common setup connects speech recognition, a language model and speech synthesis. It also needs turn detection, playback, interruption handling and any application tools. A realtime transcription model only handles the recognition stage; a streaming speech model only handles the spoken output. Test the complete conversation before choosing based on one stage’s speed.

For a finished video or podcast

Lock the script, render a voiceover, generate or choose a music bed, then edit and mix. Review voice permissions and each service or model license. Keep a clean narration track for captions, language versions and future revisions.

Pair your audio with a video model →

Explore model and platform availability →