xAI Grok TTS
TTS naturel et expressif — #3 Humanness Index Vapi (93/100), 460ms TTFA, 5 voix, 20 langues, $15/1M chars
Comparative Scores
Architecture
Medium relevance: #3 Humanness Index and 285ms TTFA make it competitive for Storygami/Edugami voice agents. However, US-only hosting and no GDPR documentation are blockers for Swiss/EU deployment. No lip-sync timestamps limits avatar integration. Best evaluated as part of the Grok Voice Think Fast 2.0 S2S pipeline ($0.08/min all-in) rather than standalone TTS.
Analysis
Grok TTS is the text-to-speech model behind Grok Voice, Tesla vehicles, and Starlink customer support. It ranks #3 on the Vapi Humanness Index (93/100, 1050 blind votes) — behind Speechify Simba 3.2 (#1) and ElevenLabs Eleven v3 (#2) — with 460ms median TTFA (285ms with optimize_streaming_latency). Five expressive voices (eve, ara, rex, sal, leo) across 20 languages, with fine-grained delivery control via inline speech tags ([laugh], [whisper], <emphasis>, <pause>). Custom voice cloning via short reference clip. Part of the Grok Voice Think Fast 2.0 S2S pipeline at $0.08/min all-in.
Strengths
- #3 Humanness Index Vapi (93/100, 1050 votes) — très naturel
- 285ms TTFA (streaming optimisé)
- Speech tags inline expressifs ([laugh], [whisper], <emphasis>)
- Clonage vocal via Custom Voices API
- 5 voix intégrées, 20 langues
- Même stack que Grok Voice Think Fast 2.0 (pipeline S2S $0.08/min)
- Pronunciation replacements pour contrôle phonétique
Weaknesses
- Cloud US uniquement — pas de région EU, RGPD à vérifier
- Pas d'option on-premise
- Pas de timestamps/visèmes pour lip-sync avatar
- 5 voix intégrées seulement (vs 3000+ ElevenLabs)
- Pay-as-you-go uniquement — pas de plan d'abonnement
- Pas de score ELO Artificial Analysis TTS
Voice Capabilities
Clonage vocal via Custom Voices API (clip de référence court). Le voice_id résultant fonctionne comme une voix intégrée dans session.update.
Style tags and non-verbals allow breath, intensity, speed and intent to be prescribed in text.
How: `voice_id`, `speed`, `replace`, `[pause]`, `[laugh]`, `<whisper>`, `<slow>`, `<build-intensity>` and other inline tags.
Validate: Test each tag with target voice and language; styles remain generative interpretation.
WebSocket streaming temps réel (optimize_streaming_latency : 285ms TTFA). REST API pour batch. Speech tags inline : [laugh], [sigh], [whisper], <emphasis>, <slow>, <pause>. Vitesse de parole ajustable (0.7–1.5×).
Pas de timestamps/visèmes natifs pour lip-sync. À combiner avec un STT pour obtenir des timestamps.
Pricing
$15.00/1M chars (~$0.0075/min estimé). Pay-as-you-go uniquement. Pas de plan d'abonnement.
| Plan | Subscription/mo | Included | Overage/min |
|---|---|---|---|
Pay-as-you-goRecommended $15.00/1M chars (~$0.0075/min) | Free | Pay-as-you-go | $0.0075 |
Sovereignty & Compliance
Cloud uniquement (US). Pas d'option on-premise documentée.
Data residency: US (xAI). Pas de région EU documentée. RGPD à vérifier.
This sheet's verification
Vapi Humanness Index #3 (93/100, 1050 votes, 30 juil. 2026). TTFA 460ms médian / 285ms streaming optimisé (Vapi benchmark, juin 2026).
Update note: Ajout initial — lancement standalone STT/TTS API xAI (17 avril 2026). Grok Voice Think Fast 2.0 lancé le 29 juillet 2026. Humanness Index #3 (30 juil. 2026).
This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.