TTS & Voice Synthesis
Comparison of voice synthesis solutions for conversational pipelines (2025–2026). Benchmarks, strategic stakes, and decision questions.
Strategic framing: hear TTS in context
What the numbers say
TTFA measures audible start; ELO is a preference signal, not a guarantee for your script. Treat both as signals, then listen to short turns, pauses, numbers and sensitive lines.
What shapes the experience
Cloning, languages and especially performance direction determine how the voice speaks. Distinguish explicit control from simple instructions: test whether intent, pace and pauses hold from one turn to the next.
Stay alert to
Price per hour is an entry point, not a production budget: include volume, streams, caching and regeneration. Also check cloning consent and rights, data residency, and fallback if a voice or API changes.
Reading method: set the required control and voice level first, then compare latency and quality on the same lines, and finally test cost, rights and fallback across a complete session.
This section compares cloud streaming TTS APIs (2025–2026). They speed up integration, with vendor dependency and sovereignty constraints to assess for each deployment context.
Click a header to sort
| Solution | TTFA | ELO | Cloning | Performance direction | Multilingual | Price/hr · entry | Detail |
|---|---|---|---|---|---|---|---|
StepAudio 2.5 TTS Benchmarking qualité — non recommandé pour production EU | 200 ms | 1205 | ✓ | 7/10 · Instruction-based direction | ✓ 100 | $4.59/h | View → |
Inworld Realtime TTS-2 + FlashUpd Phase 1 MVP — Qualité + Pipeline conversationnel complet + 100+ langues | 100 ms | 1198 | ✓ | 7/10 · Instruction-based direction | ✓ 200 | $1.35/h | View → |
Eleven v4 / v4 TurboUpd Phase 1 MVP — Référence qualité | 150 ms | 1177 | ✓ | 7/10 · Instruction-based direction | ✓ 90 | $12.00/h | View → |
MiniMax Speech 2.8New Prototypage économique — non recommandé pour production EU | 200 ms | 1175 | ✓ | 10/10 · Explicit API control | ✓ 17 | $3.24/h | View → |
Fish Audio OpenAudio S1 Phase 1 MVP — Coût/Souveraineté | 200 ms | 1124 | ✓ | 7/10 · Instruction-based direction | ✓ 13 | $0.81/h | View → |
Gradium TTSNew Phase B R&D — Lip-sync + Infrastructure Temps Réel | 216 ms | 1096 | ✓ | 3/10 · Limited control | ✓ 5 | $2.60/h | View → |
Cartesia Sonic 3.6Upd Phase 1 MVP — Latence critique | 90 ms | 1072 | ✓ | 10/10 · Explicit API control | ✓ 44 | $2.26/h | View → |
OpenAI Realtime (gpt-realtime-2.1) + GPT-Live-1Upd Phase 1 MVP — Référence benchmark + Traduction multilingue | 300 ms | 1070 | ✗ | 7/10 · Instruction-based direction | ✓ 70 | $4.61/h | View → |
Hume AI Octave 2 Phase 1 MVP — Expressivité émotionnelle | 100 ms | 1057 | ✓ | 10/10 · Explicit API control | ✓ 11 | $8.10/h | View → |
Smallest.ai Lightning V3.1 Pipeline complète — Phase B Hydra S2S | 100 ms | 1024 | ✓ | 3/10 · Limited control | ✓ 15 | $0.94/h | View → |
Gemini 3.8 Flash TTSNew Direction de jeu vocale — à tester en français | unpublished | — | ✓ | 7/10 · Instruction-based direction | ✓ 130 | API price pending | View → |
Gemini 3.8 Flash-Lite TTSNew Cascade temps réel — à chiffrer | unpublished | — | ✓ | 7/10 · Instruction-based direction | ✓ 101 | API price pending | View → |
Deepgram Flux TTSNew Phase B R&D — Test conversationnel temps réel | 80 ms | — | ✗ | 10/10 · Explicit API control | ✗ | $2.43/h | View → |
Deepgram Aura 2Upd Phase 1 MVP — Stack ASR+TTS intégré | 80 ms | — | ✗ | 3/10 · Limited control | ✗ | $1.62/h | View → |
xAI Grok TTSNew Phase B R&D — Évaluation pipeline S2S | 285 ms | — | ✓ | 7/10 · Instruction-based direction | ✓ 20 | $0.81/h | View → |
Compare by solution
StepAudio 2.5 TTS
Contextual TTS — ELO 1187, dual-level context control, zero-shot voice cloning in 3 sec
Instruction-based direction
Inworld Realtime TTS-2 + Flash
TTS conversationnel disponible : direction libre, contexte audio, non-verbaux et variante Flash à faible délai serveur
Instruction-based direction
Eleven v4 / v4 Turbo
Eleven v4 (28 sept. 2026) — 90+ langues, tags audio et Turbo ~150 ms premier son (éditeur). v3 reste le repère ELO d’août.
Instruction-based direction
MiniMax Speech 2.8
Sound tags natifs — $0.10/1M chars, 17 langues, clonage vocal instant
Explicit API control
Fish Audio OpenAudio S1
Pay-as-you-go voice cloning — 70% cheaper than ElevenLabs
Instruction-based direction
Gradium TTS
TTS streaming français avec précision de prononciation structurée et résidence UE activable
Limited control
Cartesia Sonic 3.6
Sonic 3.6 (27 août 2026) — 44 langues, plafond éditeur sous 90 ms. ELO fiche = snapshot du 18 août.
Explicit API control
OpenAI Realtime (gpt-realtime-2.1) + GPT-Live-1
2.1 : interruptions et bruit annoncés améliorés. GPT-Live-1 (10 sept.) : full-duplex à 0,05 $/min, plus le modèle derrière.
Instruction-based direction
Hume AI Octave 2
LLM-based emotional TTS — natural language emotion control
Explicit API control
Smallest.ai Lightning V3.1
Full voice pipeline (TTS + STT + LLM + S2S) — sub-100ms, 15+ languages
Limited control
Gemini 3.8 Flash TTS
Nouveau — design de voix et acting ligne à ligne, 130 langues. Pas de TTFA publié.
Instruction-based direction
Gemini 3.8 Flash-Lite TTS
Nouveau — variante débit et agents en cascade, 101 langues. Pas de TTFA publié.
Instruction-based direction
Deepgram Flux TTS
New — conversation-native TTS: 80ms first audio (vendor), cross-turn context, WebSocket + REST
Explicit API control
Deepgram Aura 2
Aura 2 baseline — Coval: 327ms median TTFA, 5.1% WER (12 Aug 2026)
Limited control
xAI Grok TTS
TTS naturel et expressif — #3 Humanness Index Vapi (93/100), 460ms TTFA, 5 voix, 20 langues, $15/1M chars
Instruction-based direction
Global architecture: Voice-to-Voice & real time
Cascade, speech-to-speech and full-duplex comparison lives in a separate journey so this page remains focused on voice synthesis.
Explore Voice-to-Voice →