GamiWays

Custom Voice Tool Ranking

Weight criteria to your context and get a dynamic TTS / STT ranking.

Preset Profiles

Criteria Weights

0–10
Voice Quality5
Latency (TTFA)5
Voice Cloning5
Expressiveness / Direction5
Data Sovereignty5
Cost / Pricing5
Multilingual5

Each criterion is scored 0–10. A weight of 0 excludes the criterion from the calculation. The final score is the weighted average.

TTS Ranking — 25 tools

Sorted by weighted score
🥇
Cloud APIUpd

Inworld Realtime TTS-2 + Flash

TTS conversationnel disponible : direction libre, contexte audio, non-verbaux et variante Flash à faible délai serveur

Sov. mediumLock-in medium
8.4
/10
Quality×5
9
Latency×5
8
Cloning×5
9
Expressiveness / Direction×5
10
Sovereignty×5
6
Pricing×5
7
100ms TTFAELO 1198Commercial (training framework open-sourced)
🥈
Open Source

Miso-TTS v1 (8B)

110ms latency — most emotive open-source TTS with RVQ architecture

8.3
/10
Quality×5
9
Latency×5
9
Cloning×5
9
Expressiveness / Direction×5
10
Sovereignty×5
10
Pricing×5
10
110ms TTFAModified MIT (open-source weights on Hugging Face)
🥉
Open Source

Orpheus 3B

LLM-based TTS — ultra-natural speech with emotion tags and non-verbals

Sov. highLock-in low
7.9
/10
Quality×5
8
Latency×5
6
Cloning×5
7
Expressiveness / Direction×5
9
Sovereignty×5
10
Pricing×5
10
200ms TTFAApache 2.0
4
Cloud APIUpd

Eleven v4 / v4 Turbo

Eleven v4 (28 sept. 2026) — 90+ langues, tags audio et Turbo ~150 ms premier son (éditeur). v3 reste le repère ELO d’août.

Sov. mediumLock-in medium
7.1
/10
Quality×5
8
Latency×5
8
Cloning×5
10
Expressiveness / Direction×5
10
Sovereignty×5
2
Pricing×5
2
150ms TTFAELO 1177Commercial
5
Open Source

Sesame CSM

Conversational Speech Model — crosses the uncanny valley of voice

Sov. highLock-in low
7.1
/10
Quality×5
9
Latency×5
3
Cloning×5
7
Expressiveness / Direction×5
10
Sovereignty×5
10
Pricing×5
10
400ms TTFAApache 2.0 (research)
6
Open Source

Dia (Nari Labs)

Ultra-realistic dialogue generation — multi-speaker, emotion, non-verbals

Sov. highLock-in low
7.0
/10
Quality×5
8
Latency×5
4
Cloning×5
7
Expressiveness / Direction×5
9
Sovereignty×5
10
Pricing×5
10
300ms TTFAApache 2.0
7
Open Source

Kyutai TTS 1.6B

Delayed streams modeling — streaming-native, timestamps, batching

Sov. highLock-in low
7.0
/10
Quality×5
7
Latency×5
8
Cloning×5
6
Expressiveness / Direction×5
6
Sovereignty×5
10
Pricing×5
10
100ms TTFACC-BY 4.0
8
Open SourceUpd

Voxtral TTS (Mistral)

Open-weights TTS from Mistral — fast, adaptable, 9 languages (Mar 2026)

Sov. highLock-in low
7.0
/10
Quality×5
4
Latency×5
8
Cloning×5
7
Expressiveness / Direction×5
6
Sovereignty×5
9
Pricing×5
9
150ms TTFAELO 1081Open weights (Mistral license)
9
Cloud API

StepAudio 2.5 TTS

Contextual TTS — ELO 1187, dual-level context control, zero-shot voice cloning in 3 sec

7.0
/10
Quality×5
9
Latency×5
6
Cloning×5
9
Expressiveness / Direction×5
10
Sovereignty×5
1
Pricing×5
5
200ms TTFAELO 1205Proprietary (StepFun / 阶跃星辰)
10
Cloud APINew

MiniMax Speech 2.8

Sound tags natifs — $0.10/1M chars, 17 langues, clonage vocal instant

7.0
/10
Quality×5
8
Latency×5
6
Cloning×5
8
Expressiveness / Direction×5
9
Sovereignty×5
1
Pricing×5
10
200ms TTFAELO 1175Commercial
11
Cloud APINew

xAI Grok TTS

TTS naturel et expressif — #3 Humanness Index Vapi (93/100), 460ms TTFA, 5 voix, 20 langues, $15/1M chars

6.7
/10
Quality×5
8
Latency×5
8
Cloning×5
7
Expressiveness / Direction×5
8
Sovereignty×5
2
Pricing×5
7
285ms TTFACommercial
12
Cloud APIUpd

Cartesia Sonic 3.6

Sonic 3.6 (27 août 2026) — 44 langues, plafond éditeur sous 90 ms. ELO fiche = snapshot du 18 août.

Sov. lowLock-in medium
6.6
/10
Quality×5
4
Latency×5
9
Cloning×5
8
Expressiveness / Direction×5
8
Sovereignty×5
4
Pricing×5
5
90ms TTFAELO 1072Commercial
13
Cloud APINew

Gemini 3.8 Flash TTS

Nouveau — design de voix et acting ligne à ligne, 130 langues. Pas de TTFA publié.

6.4
/10
Quality×5
8
Latency×5
4
Cloning×5
7
Expressiveness / Direction×5
9
Sovereignty×5
2
Pricing×5
5
0ms TTFACommercial
14
Open Source

Chatterbox (Resemble AI)

MIT license — beats ElevenLabs in blind tests (63.75% preference)

Sov. highLock-in low
6.4
/10
Quality×5
2
Latency×5
7
Cloning×5
8
Expressiveness / Direction×5
8
Sovereignty×5
10
Pricing×5
9
150ms TTFAELO 1021MIT
15
Cloud API

Smallest.ai Lightning V3.1

Full voice pipeline (TTS + STT + LLM + S2S) — sub-100ms, 15+ languages

6.4
/10
Quality×5
2
Latency×5
9
Cloning×5
7
Expressiveness / Direction×5
6
Sovereignty×5
6
Pricing×5
8
100ms TTFAELO 1024Commercial
16
Cloud APINew

Gradium TTS

TTS streaming français avec précision de prononciation structurée et résidence UE activable

6.4
/10
Quality×5
5
Latency×5
7
Cloning×5
8
Expressiveness / Direction×5
3
Sovereignty×5
6
Pricing×5
9
216ms TTFAELO 1096Commercial
17
Cloud APINew

Gemini 3.8 Flash-Lite TTS

Nouveau — variante débit et agents en cascade, 101 langues. Pas de TTFA publié.

6.1
/10
Quality×5
7
Latency×5
5
Cloning×5
6
Expressiveness / Direction×5
8
Sovereignty×5
2
Pricing×5
6
0ms TTFACommercial
18
Cloud API

Fish Audio OpenAudio S1

Pay-as-you-go voice cloning — 70% cheaper than ElevenLabs

Sov. highLock-in low
6.1
/10
Quality×5
6
Latency×5
6
Cloning×5
8
Expressiveness / Direction×5
7
Sovereignty×5
4
Pricing×5
7
200ms TTFAELO 1124Commercial
19
Open Source

Ultravox v0.5

Speech-to-speech model — ~100ms latency, no ASR/TTS pipeline needed

Sov. highLock-in low
6.1
/10
Quality×5
7
Latency×5
10
Cloning×5
1
Expressiveness / Direction×5
6
Sovereignty×5
7
Pricing×5
7
100ms TTFACommercial API (CC-BY-NC-4.0 weights)
20
Open Source

Moshi (Kyutai)

Full-duplex spoken dialogue — simultaneous listening and speaking

Sov. highLock-in low
6.1
/10
Quality×5
7
Latency×5
8
Cloning×5
1
Expressiveness / Direction×5
6
Sovereignty×5
10
Pricing×5
10
200ms TTFACC-BY 4.0
21
Cloud API

Hume AI Octave 2

LLM-based emotional TTS — natural language emotion control

Sov. lowLock-in high
5.9
/10
Quality×5
3
Latency×5
8
Cloning×5
6
Expressiveness / Direction×5
10
Sovereignty×5
2
Pricing×5
8
100ms TTFAELO 1057Commercial
22
Open Source

Kokoro 82M v1.0

Highest-ranked open-weight TTS — ELO 1059, 82M params, Apache 2.0

Sov. highLock-in low
5.7
/10
Quality×5
3
Latency×5
9
Cloning×5
1
Expressiveness / Direction×5
5
Sovereignty×5
10
Pricing×5
10
60ms TTFAELO 1059Apache 2.0
23
Cloud APINew

Deepgram Flux TTS

New — conversation-native TTS: 80ms first audio (vendor), cross-turn context, WebSocket + REST

Sov. mediumLock-in high
5.1
/10
Quality×5
7
Latency×5
9
Cloning×5
1
Expressiveness / Direction×5
8
Sovereignty×5
6
Pricing×5
4
80ms TTFACommercial
24
Cloud APIUpd

OpenAI Realtime (gpt-realtime-2.1) + GPT-Live-1

2.1 : interruptions et bruit annoncés améliorés. GPT-Live-1 (10 sept.) : full-duplex à 0,05 $/min, plus le modèle derrière.

Sov. lowLock-in high
4.7
/10
Quality×5
4
Latency×5
6
Cloning×5
1
Expressiveness / Direction×5
8
Sovereignty×5
2
Pricing×5
3
300ms TTFAELO 1070Commercial
25
Cloud APIUpd

Deepgram Aura 2

Aura 2 baseline — Coval: 327ms median TTFA, 5.1% WER (12 Aug 2026)

Sov. mediumLock-in medium
4.4
/10
Quality×5
6
Latency×5
9
Cloning×5
1
Expressiveness / Direction×5
4
Sovereignty×5
3
Pricing×5
7
80ms TTFACommercial

Methodology: Raw scores (1–10) are sourced from public benchmarks (Artificial Analysis ELO, Koenecke WER, measured TTFA). Weighting is applied via weighted average. Sovereignty and lock-in badges come from the GamiWays strategic analysis.