GamiWays

TTS & Voice Synthesis

Comparison of voice synthesis solutions for conversational pipelines (2025–2026). Benchmarks, strategic stakes, and decision questions.

🎯

Strategic framing: hear TTS in context

What the numbers say

TTFA measures audible start; ELO is a preference signal, not a guarantee for your script. Treat both as signals, then listen to short turns, pauses, numbers and sensitive lines.

What shapes the experience

Cloning, languages and especially performance direction determine how the voice speaks. Distinguish explicit control from simple instructions: test whether intent, pace and pauses hold from one turn to the next.

Stay alert to

Price per hour is an entry point, not a production budget: include volume, streams, caching and regeneration. Also check cloning consent and rights, data residency, and fallback if a voice or API changes.

Reading method: set the required control and voice level first, then compare latency and quality on the same lines, and finally test cost, rights and fallback across a complete session.

CLOUD APIs

This section compares cloud streaming TTS APIs (2025–2026). They speed up integration, with vendor dependency and sovereignty constraints to assess for each deployment context.

Click a header to sort

SolutionTTFAELOCloningPerformance directionMultilingualPrice/hr · entryDetail
StepAudio 2.5 TTS
Benchmarking qualité — non recommandé pour production EU
200 ms1205✓7/10 · Instruction-based direction✓ 100$4.59/hView →
Inworld Realtime TTS-2 + FlashUpd
Phase 1 MVP — Qualité + Pipeline conversationnel complet + 100+ langues
100 ms1198✓7/10 · Instruction-based direction✓ 200$1.35/hView →
Eleven v4 / v4 TurboUpd
Phase 1 MVP — Référence qualité
150 ms1177✓7/10 · Instruction-based direction✓ 90$12.00/hView →
MiniMax Speech 2.8New
Prototypage économique — non recommandé pour production EU
200 ms1175✓10/10 · Explicit API control✓ 17$3.24/hView →
Fish Audio OpenAudio S1
Phase 1 MVP — Coût/Souveraineté
200 ms1124✓7/10 · Instruction-based direction✓ 13$0.81/hView →
Gradium TTSNew
Phase B R&D — Lip-sync + Infrastructure Temps Réel
216 ms1096✓3/10 · Limited control✓ 5$2.60/hView →
Cartesia Sonic 3.6Upd
Phase 1 MVP — Latence critique
90 ms1072✓10/10 · Explicit API control✓ 44$2.26/hView →
OpenAI Realtime (gpt-realtime-2.1) + GPT-Live-1Upd
Phase 1 MVP — Référence benchmark + Traduction multilingue
300 ms1070✗7/10 · Instruction-based direction✓ 70$4.61/hView →
Hume AI Octave 2
Phase 1 MVP — Expressivité émotionnelle
100 ms1057✓10/10 · Explicit API control✓ 11$8.10/hView →
Smallest.ai Lightning V3.1
Pipeline complète — Phase B Hydra S2S
100 ms1024✓3/10 · Limited control✓ 15$0.94/hView →
Gemini 3.8 Flash TTSNew
Direction de jeu vocale — à tester en français
unpublished—✓7/10 · Instruction-based direction✓ 130API price pendingView →
Gemini 3.8 Flash-Lite TTSNew
Cascade temps réel — à chiffrer
unpublished—✓7/10 · Instruction-based direction✓ 101API price pendingView →
Deepgram Flux TTSNew
Phase B R&D — Test conversationnel temps réel
80 ms—✗10/10 · Explicit API control✗$2.43/hView →
Deepgram Aura 2Upd
Phase 1 MVP — Stack ASR+TTS intégré
80 ms—✗3/10 · Limited control✗$1.62/hView →
xAI Grok TTSNew
Phase B R&D — Évaluation pipeline S2S
285 ms—✓7/10 · Instruction-based direction✓ 20$0.81/hView →

Compare by solution

Cloud API

StepAudio 2.5 TTS

Contextual TTS — ELO 1187, dual-level context control, zero-shot voice cloning in 3 sec

Instruction-based direction

Quality9/10
9
Latency6/10
6
Cloning9/10
9
Performance direction7/10
7
Sovereignty1/10
1
Pricing5/10
5
200 ms TTFAELO 1205Cloning
Benchmarking qualité — non recommandé pour production EU
Full details →
Cloud APIUpdated

Inworld Realtime TTS-2 + Flash

TTS conversationnel disponible : direction libre, contexte audio, non-verbaux et variante Flash à faible délai serveur

Instruction-based direction

Quality9/10
9
Latency8/10
8
Cloning9/10
9
Performance direction7/10
7
Sovereignty6/10
6
Pricing7/10
7
100 ms TTFAELO 1198CloningSovereignLip-sync
Phase 1 MVP — Qualité + Pipeline conversationnel complet + 100+ langues
Full details →
Cloud APIUpdated

Eleven v4 / v4 Turbo

Eleven v4 (28 sept. 2026) — 90+ langues, tags audio et Turbo ~150 ms premier son (éditeur). v3 reste le repère ELO d’août.

Instruction-based direction

Quality8/10
8
Latency8/10
8
Cloning10/10
10
Performance direction7/10
7
Sovereignty2/10
2
Pricing2/10
2
150 ms TTFAELO 1177CloningLip-sync
Phase 1 MVP — Référence qualité
Full details →
Cloud APINew

MiniMax Speech 2.8

Sound tags natifs — $0.10/1M chars, 17 langues, clonage vocal instant

Explicit API control

Quality8/10
8
Latency6/10
6
Cloning8/10
8
Performance direction10/10
10
Sovereignty1/10
1
Pricing10/10
10
200 ms TTFAELO 1175Cloning
Prototypage économique — non recommandé pour production EU
Full details →
Cloud API

Fish Audio OpenAudio S1

Pay-as-you-go voice cloning — 70% cheaper than ElevenLabs

Instruction-based direction

Quality6/10
6
Latency6/10
6
Cloning8/10
8
Performance direction7/10
7
Sovereignty4/10
4
Pricing7/10
7
200 ms TTFAELO 1124Cloning
Phase 1 MVP — Coût/Souveraineté
Full details →
Cloud APINew

Gradium TTS

TTS streaming français avec précision de prononciation structurée et résidence UE activable

Limited control

Quality5/10
5
Latency7/10
7
Cloning8/10
8
Performance direction3/10
3
Sovereignty6/10
6
Pricing9/10
9
216 ms TTFAELO 1096CloningSovereignLip-sync
Phase B R&D — Lip-sync + Infrastructure Temps Réel
Full details →
Cloud APIUpdated

Cartesia Sonic 3.6

Sonic 3.6 (27 août 2026) — 44 langues, plafond éditeur sous 90 ms. ELO fiche = snapshot du 18 août.

Explicit API control

Quality4/10
4
Latency9/10
9
Cloning8/10
8
Performance direction10/10
10
Sovereignty4/10
4
Pricing5/10
5
90 ms TTFAELO 1072Cloning
Phase 1 MVP — Latence critique
Full details →
Cloud APIUpdated

OpenAI Realtime (gpt-realtime-2.1) + GPT-Live-1

2.1 : interruptions et bruit annoncés améliorés. GPT-Live-1 (10 sept.) : full-duplex à 0,05 $/min, plus le modèle derrière.

Instruction-based direction

Quality4/10
4
Latency6/10
6
Cloning1/10
1
Performance direction7/10
7
Sovereignty2/10
2
Pricing3/10
3
300 ms TTFAELO 1070
Phase 1 MVP — Référence benchmark + Traduction multilingue
Full details →
Cloud API

Hume AI Octave 2

LLM-based emotional TTS — natural language emotion control

Explicit API control

Quality3/10
3
Latency8/10
8
Cloning6/10
6
Performance direction10/10
10
Sovereignty2/10
2
Pricing8/10
8
100 ms TTFAELO 1057Cloning
Phase 1 MVP — Expressivité émotionnelle
Full details →
Cloud API

Smallest.ai Lightning V3.1

Full voice pipeline (TTS + STT + LLM + S2S) — sub-100ms, 15+ languages

Limited control

Quality2/10
2
Latency9/10
9
Cloning7/10
7
Performance direction3/10
3
Sovereignty6/10
6
Pricing8/10
8
100 ms TTFAELO 1024CloningSovereign
Pipeline complète — Phase B Hydra S2S
Full details →
Cloud APINew

Gemini 3.8 Flash TTS

Nouveau — design de voix et acting ligne à ligne, 130 langues. Pas de TTFA publié.

Instruction-based direction

Quality8/10
8
Latency4/10
4
Cloning7/10
7
Performance direction7/10
7
Sovereignty2/10
2
Pricing5/10
5
unpublished TTFACloning
Direction de jeu vocale — à tester en français
Full details →
Cloud APINew

Gemini 3.8 Flash-Lite TTS

Nouveau — variante débit et agents en cascade, 101 langues. Pas de TTFA publié.

Instruction-based direction

Quality7/10
7
Latency5/10
5
Cloning6/10
6
Performance direction7/10
7
Sovereignty2/10
2
Pricing6/10
6
unpublished TTFACloning
Cascade temps réel — à chiffrer
Full details →
Cloud APINew

Deepgram Flux TTS

New — conversation-native TTS: 80ms first audio (vendor), cross-turn context, WebSocket + REST

Explicit API control

Quality7/10
7
Latency9/10
9
Cloning1/10
1
Performance direction10/10
10
Sovereignty6/10
6
Pricing4/10
4
80 ms TTFASovereign
Phase B R&D — Test conversationnel temps réel
Full details →
Cloud APIUpdated

Deepgram Aura 2

Aura 2 baseline — Coval: 327ms median TTFA, 5.1% WER (12 Aug 2026)

Limited control

Quality6/10
6
Latency9/10
9
Cloning1/10
1
Performance direction3/10
3
Sovereignty3/10
3
Pricing7/10
7
80 ms TTFA
Phase 1 MVP — Stack ASR+TTS intégré
Full details →
Cloud APINew

xAI Grok TTS

TTS naturel et expressif — #3 Humanness Index Vapi (93/100), 460ms TTFA, 5 voix, 20 langues, $15/1M chars

Instruction-based direction

Quality8/10
8
Latency8/10
8
Cloning7/10
7
Performance direction7/10
7
Sovereignty2/10
2
Pricing7/10
7
285 ms TTFACloning
Phase B R&D — Évaluation pipeline S2S
Full details →

Global architecture: Voice-to-Voice & real time

Cascade, speech-to-speech and full-duplex comparison lives in a separate journey so this page remains focused on voice synthesis.

Explore Voice-to-Voice →