Text becomes a finished voiceover. Delivery, pronunciation and consistency matter most.
Voice & audio · September 2026
Give your project the right voice.
Narration, live conversation, transcription and music are different jobs. Match the model to the output you need, then listen to it on a real script.
Speech arrives while someone interacts. Measure pauses, interruptions and time to audible response.
Recorded or live speech becomes text, timestamps and speaker labels.
A description or lyrics become a song or soundtrack. Arrange and mix separately from narration.
Text to speech
Six speech choices worth testing.
Choose two candidates for your specific language, voice and delivery. Provider latency claims do not measure the complete experience in your app.
MiniMax Speech 2.8 HD
Try for polished narration and multilingual voiceovers.
Where to use it
MiniMax API
Text-to-speech with streaming, pronunciation controls and voice selection. Compare HD and Turbo on the same script.
Model identifier
speech-2.8-hdMiniMax Speech 2.8 Turbo
Try for interactive or repeated speech generation.
Where to use it
MiniMax API
A separate speech tier. Measure response time, pronunciation and delivered-audio cost on your workload.
Model identifier
speech-2.8-turboCartesia Sonic 3.6
Test conversational speech and responsive spoken interfaces.
Where to use it
Cartesia API / Playground
Generally available August 27, 2026; 44 languages. Pin sonic-3.6-2026-08-27 for a reproducible model snapshot.
Model identifier
sonic-3.6Eleven v3
Expressive narration and multi-speaker dialogue.
Where to use it
ElevenLabs API / creative tools
Choose for performance and delivery; evaluate the actual voice and script.
Model identifier
eleven_v3Eleven v3 Conversational
Expressive realtime speech synthesis.
Where to use it
ElevenLabs API
Supports realtime dialogue through a Text to Dialogue WebSocket.
Model identifier
eleven_v3_conversationalEleven Flash v2.5
A speed-focused speech option for interactive applications.
Where to use it
ElevenLabs API
Test dates, currencies and phone numbers: text normalization differs from other models.
Model identifier
eleven_flash_v2_5Speech to text
Turn the recording into usable text.
Test names, accents, overlapping speakers and background noise before relying on an automatic transcript.
Scribe v2
Transcribe recordings with speaker labels and word timestamps.
Where to use it
ElevenLabs API
Speech-to-text, rather than a voice generator.
Model identifier
scribe_v2Scribe v2 Realtime
Streaming transcription for live audio.
Where to use it
ElevenLabs API
Evaluate transcript revisions, silence handling and noisy speech.
Model identifier
scribe_v2_realtimeMusic generation
Build the soundtrack separately.
Keep dialogue, music and effects on separate tracks so a change to one does not force you to regenerate everything.
MiniMax Music 3
Turn a musical brief and lyrics into a complete song.
Where to use it
Downloadable weights / ComfyUI
Up to five minutes, 32 kHz stereo output. Review the model’s license and official workflow requirements before production use.
Model identifier
MiniMaxAI/MiniMax-Music3Choose a first voice test
For a finished voiceover, compare MiniMax Speech 2.8 HD with Eleven v3. For an interactive interface, compare Sonic 3.6, Eleven v3 Conversational, Flash v2.5 or Speech 2.8 Turbo. These are suggested experiments, not a measured ranking.
Use a short script containing a person’s name, a company name, a date, a price, an acronym and a sentence that changes emotion. Include your actual target language. First record how each should sound, then listen without looking at which model produced the result.
- Check pronunciation. Review names, abbreviations, dates and amounts. Save corrections with the script.
- Check performance. Listen for inappropriate emphasis, pauses, emotional delivery and consistency across a longer passage.
- Measure the whole interaction. For live use, include recognition, reasoning, network delays, first audible speech and interruption handling.
- Track usable output cost. Count retries, speech minutes, credits or characters, voice-plan fees and editing time. Compare the same final recording length.
A speech model is one part of a voice assistant
A common setup connects speech recognition, a language model and speech synthesis. It also needs turn detection, playback, interruption handling and any application tools. A realtime transcription model only handles the recognition stage; a streaming speech model only handles the spoken output. Test the complete conversation before choosing based on one stage’s speed.
For a finished video or podcast
Lock the script, render a voiceover, generate or choose a music bed, then edit and mix. Review voice permissions and each service or model license. Keep a clean narration track for captions, language versions and future revisions.