Deepgram Flux TTS
New — conversation-native TTS: 80ms first audio (vendor), cross-turn context, WebSocket + REST
Comparative Scores
Architecture
English-only conversation-native TTS. Useful to study interruption reconciliation, not as the French voice for Où est Ava ?. The free launch offer ended on 12 September 2026. Compare later with Sonic 3.6, Inworld TTS-2 and Gradium only if a French voice appears.
Analysis
Deepgram Flux TTS (GA, 12 August 2026) is a conversation-native TTS family for real-time voice agents. Rather than restarting each response, it retains cross-turn context, reports what the caller actually heard after an interruption, and lets the agent adjust speed, pacing and pronunciation during speech. Deepgram reports first audio as low as 80ms and under 200ms independently of response length. Its internal read-aloud benchmark claims 2.2% median WER and 3.4% on hard prompts, but no public independent ranking has yet been identified for Flux TTS. In contrast, Coval/Openbenchmarks currently measures the older Aura-2-Thalia-en at 327ms median TTFA, 547ms p95 and 5.1% WER; these independent Aura results must not be attributed to Flux.
Strengths
- Conversation-native: cross-turn context, lifecycle events and interruption reconciliation
- 80ms first-audio / <200ms response-length-independent latency claimed by Deepgram
- WebSocket streaming plus REST batch on the same /v2/speak endpoint
- Paired Flux STT + Flux TTS configuration for a unified real-time loop
- Cloud, self-hosted and on-premise announced
Weaknesses
- GA English voices only; multilingual voices are roadmap
- No voice cloning at launch
- No documented word/phoneme/viseme timestamps for avatar lip-sync
- $45/1M characters after launch offer — 3× Aura-2 list rate
- 80ms and WER claims are vendor benchmarks; no independent Flux TTS score yet
Voice Capabilities
Non disponible au lancement ; voice cloning annoncé sur la roadmap Deepgram.
Flux allows control of pace and broad conversational expressivity across turns.
How: `speed` (0.85–1.15) and beta `expressivity` (-2 to 2) on `/v2/speak`.
Validate: This is not a dictionary of per-line emotions: validate granularity and beta availability on the chosen voice model.
WebSocket streaming à /v2/speak : génération audio dès le premier token, cycle de parole, interruption, reprise et contexte inter-tours. Deepgram annonce 80ms de premier audio au mieux et <200ms quelle que soit la longueur de réponse.
Aucun timestamp mot/phonème/visème documenté au lancement.
Pricing
L’offre gratuite s’est terminée le 12 septembre 2026. Tarif ensuite : $45/1M caractères PAYG ; $40.50/1M sur Growth.
| Plan | Subscription/mo | Included | Overage/min | Top-up |
|---|---|---|---|---|
Pay As You Go Standard concurrency after launch is plan/contract dependent — confirm with Deepgram. | Free | $45 / 1M chars | $0.0450 | Yes |
GrowthRecommended 10% discounted character rate; prepaid Growth credits and concurrency terms to confirm. | Free | $40.50 / 1M chars | $0.0405 | Yes |
Sovereignty & Compliance
Cloud, self-hosted et on-premise annoncés ; confirmer les modalités, la résidence UE et les garanties contractuelles avec Deepgram.
Data residency: Régions cloud non précisées publiquement pour Flux TTS ; self-hosted/on-premise annoncés, à qualifier pour Suisse/UE.
Deepgram Flux TTS — Strategic Positioning
Beyond technical specs: where does this tool sit in the ecosystem, what are the risks and strategic implications for GamiWays?
Flux TTS shifts Deepgram's offer from fast narration to a stateful conversational speech layer. Its promise is operational simplicity, not a verified independent quality lead — evaluate the vendor claims under Où est Ava ?-like multi-turn load before committing.
A. Strategic Positioning
Target customer: Voice-agent developer / Enterprise — English real-time conversational agents
Conversation-native TTS: cross-turn context, interruption reconciliation and streaming-first speech generation are provided by the model rather than assembled through client-side orchestration.
B. Competitive Moat
- Flux STT + Flux TTS paired stack: one vendor, one configuration and an explicit speech lifecycle
- Mamba state-space architecture and interleaved text/audio generation target low latency without resetting session context
- Cloud, self-hosted and on-premise deployment announced for regulated voice-agent workloads
Vulnerability: New English-only model with no voice cloning and no public independent Flux benchmark yet. Post-launch pricing is materially above Aura-2 and self-hosted/EU terms still require contractual verification.
E. Strategic Questions for GamiWays
Sovereignty fit
Self-hosted and on-premise are announced, but deployment evidence, EU residency and commercial terms need validation before treating Flux as a sovereign option.
Build vs. Buy
Buy as a Phase B experiment for English agents. Preserve a modular TTS adapter and keep Cartesia, Inworld or Gradium as comparable paths until Flux is independently benchmarked.
Lock-in risk
The most differentiated capabilities — conversational state, interruption reconciliation and future shared STT/TTS state — are proprietary platform behaviors that increase switching costs.
Roadmap alignment
High experimental fit for real-time conversation and interruption handling; conditional fit for Où est Ava ? until French/German, avatar alignment data and EU/on-premise deployment are validated.
This sheet's verification
Deepgram Flux TTS launch benchmark (vendor, 12 Aug 2026). Independent Flux ranking not yet available. Coval/Openbenchmarks Aura-2 baseline refreshed 12 Aug 2026.
Update note: 30 septembre 2026 : l’offre gratuite du 12 août au 12 septembre est terminée. Pas de nouveau modèle. Les voix restent anglaises.
This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.