Realtime TTS-2 and its Flash variant are available. Voice Direction, non-verbals and audio context can guide vocal delivery; French instruction following, reproducibility and complete Où est Ava ? latency still need measurement.
Inworld Realtime TTS-2 + Flash
TTS conversationnel disponible : direction libre, contexte audio, non-verbaux et variante Flash à faible délai serveur
Comparative Scores
Architecture
A candidate to test for the Où est Ava ? voice layer, not the video avatar layer. Voice Direction, non-verbals and cross-turn audio context make it a direct path for how a line is delivered. Test French directions, sensitive scenes, lip-sync with the chosen renderer and end-to-end latency. For Plastic Dilemma, treat EU/India residency as an Enterprise contractual option, not an automatic sovereign deployment.
Analysis
Inworld Realtime TTS-2 is available alongside a Flash variant. It combines natural-language Voice Direction, non-verbals, cross-turn audio context, Voice Design and one voice identity across 200+ languages. TTS-2 publishes a 100 ms P90 server TTFB and Flash 20 ms; neither is an Où est Ava ? end-to-end latency. The platform also exposes STT, a Realtime API and LLM Router. French direction following, viseme compatibility and Enterprise sovereignty terms must be validated.
Strengths
- Realtime TTS-2 disponible : direction libre et non-verbaux
- Contexte audio inter-tours et modes Expressive/Balanced/Stable
- TTS-2 Flash : 20 ms TTFB P90 annoncé côté serveur
- 200+ langues et une identité vocale crosslingue annoncées
- Clonage instantané et Voice Design selon le plan
- Plateforme : TTS + STT + Realtime S2S + LLM Router
- Timestamps utiles au lip-sync
- Résidence UE/Inde et SLA/DPA à contractualiser sur Enterprise
Weaknesses
- TTFB serveur hors réseau, LLM et vidéo
- Direction libre et non-verbaux en français à tester scène par scène
- Résidence UE/Inde et conditions on-premise limitées aux termes Enterprise
- Modèles et limites du Router évoluent fréquemment
Voice Capabilities
Instant voice cloning (free). Professional voice cloning (Growth/Enterprise add-on). Advanced Voice Design: create a voice from a prose description — no reference audio needed. Up to 5 custom voices (On-Demand), 100 (Creator), 1,000 (Developer), 3,000 (Growth). ElevenLabs migration tool available.
Available Realtime TTS-2 interprets open-ended performance directions, non-verbals and prior-turn audio context into conversational voice output.
How: Natural-language direction, non-verbal markers, cross-turn audio context and `Expressive`, `Balanced` or `Stable` modes, with voice/voice design according to plan.
Validate: Direction is a generative interpretation: measure French instruction following, reproducibility and lip-sync before an Où est Ava ? decision. Published TTFB is server-side only.
Streaming-native via WebSocket and REST. TTS-2 annonce 100 ms TTFB P90 côté serveur et Flash 20 ms P90 ; ces chiffres excluent le réseau, le LLM et la vidéo. Realtime API : WebSocket full-duplex S2S, détection de tour, tool calling et routage LLM. Les backchannels peuvent être diffusés pendant le raisonnement.
Word, character, phoneme, and viseme-level timestamps. Unity/Unreal SDKs with lipsync templates.
Pricing
Realtime TTS-2 : $25/1M chars On-Demand, $20 Creator, $17.50 Builder, $15 Developer et $12.50 Growth. TTS-2 Flash : $15 → $7/1M chars aux mêmes paliers. L’estimateur Inworld emploie ~1 000 caractères/min ; les crédits mensuels sont inclus dans les plans payants.
| Plan | Subscription/mo | Included | Overage/min | Top-up |
|---|---|---|---|---|
On-Demand (free) PAYG above 40 min at $25/1M chars | Free | 70 min | $0.0250 | Yes |
Creator Credits rollover. Overage at $20/1M chars (TTS-2) | $25/mo | $25/mo credits | $0.0200 | Yes |
GrowthRecommended Best rate: $12.50/1M chars (TTS-2) | $1500/mo | $1,500/mo credits | $0.0125 | Yes |
Sovereignty & Compliance
Résidence UE/Inde, SLA et DPA sont des options Enterprise. ZDR, HIPAA et BAA sont proposés en option Growth ; vérifier contractualisation, conservation des voix et éventuelles conditions on-premise séparément.
Data residency: US, EU, India
Certifications & Compliance
HIPAA, ZDR (Zero Data Retention) and BAA are Growth/Enterprise options. EU or India residency and any on-premise terms must be contracted separately; none means automatic sovereign deployment.
Inworld Realtime TTS-2 + Flash — Strategic Positioning
Beyond technical specs: where does this tool sit in the ecosystem, what are the risks and strategic implications for GamiWays?
Inworld Realtime TTS-2 is now available: Voice Direction and cross-turn context make it a credible way to test how Où est Ava ? speaks. The claimed 100/20ms server TTFB and the French emotional result still require an end-to-end scene test.
A. Strategic Positioning
Target customer: Enterprise / Developer — regulated industries, gaming, real-time agents
Realtime TTS-2 available for conversation: free-form Voice Direction, Conversational Awareness, non-verbals, Voice Design and a Flash variant. Enterprise conditions govern EU/India residency and on-premise.
B. Competitive Moat
- TTS 1.5 #1 Artificial Analysis Speech Arena (ahead of Google and ElevenLabs)
- TTS-2 : Voice Direction — natural language delivery instructions inline (unique in market)
- TTS-2 : Conversational Awareness — model hears prior audio turns, tone and pacing carry forward
- TTS-2 : 200+ languages/locales announced, one voice identity and crosslingual use
- Full platform: TTS + STT + Realtime API (WebSocket S2S) + LLM Router (220+ models)
Vulnerability: Published 100/20ms TTFB is server-side, not a full dialogue latency. French direction following, video-avatar sync and Enterprise/on-premise terms remain to be measured or contracted.
E. Strategic Questions for GamiWays
Sovereignty fit
Enterprise terms are needed for EU/India residency, SLA/DPA and any on-premise discussion. Treat ZDR, HIPAA and BAA as separately contracted features; do not classify standard access as sovereign.
Build vs. Buy
A high-priority buy/test option for Où est Ava ? voice direction, not a video-avatar purchase. Compare TTS-2 and Flash on the same French scenes; sovereignty requires an Enterprise contractual review.
Lock-in risk
Proprietary models and full-stack platform create some lock-in. Drop-in OpenAI Realtime API compatibility reduces switching cost. Competitive pricing reduces dependency risk.
Roadmap alignment
Strong alignment for the Où est Ava ? voice layer: cross-turn context may support continuity and Voice Direction can change a delivery without re-recording. Validate French instructions, non-verbals, visemes and end-to-end video timing before adoption.
This sheet's verification
Inworld Realtime TTS-2 announcement and documentation, 31 Aug–6 Sep 2026. Published TTFB is server-side P90; no score or Où est Ava ? latency is inferred from it.
Update note: Realtime TTS-2 et Flash vérifiés le 6 septembre 2026 : disponibilité, TTFB P90 serveur (100/20 ms), direction libre, non-verbaux, contexte audio, 200+ langues et tarifs de $25/$15 On-Demand à $12.50/$7 Growth. Les tests français et Où est Ava ? restent à réaliser.
This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.
ELO benchmarks / indices: Artificial Analysis · Last API sync : 18 August 2026