GamiWays
Available

Realtime TTS-2 and its Flash variant are available. Voice Direction, non-verbals and audio context can guide vocal delivery; French instruction following, reproducibility and complete Où est Ava ? latency still need measurement.

Cloud API#6 Artificial AnalysisCommercial (training framework open-sourced)

Inworld Realtime TTS-2 + Flash

TTS conversationnel disponible : direction libre, contexte audio, non-verbaux et variante Flash à faible délai serveur

100 ms
TTFA (best case)
100 ms
TTFA (typical)
$12.5/1M
Price per million chars
1198
ELO Score

Comparative Scores

Voice quality9/10
Latency8/10
Voice cloning9/10
Expressiveness10/10
Performance direction7/10
Sovereignty6/10
Price accessibility7/10
Multilingual10/10

Architecture

ArchitectureSpeechLM (streaming-native, quantization-aware) — TTS-2 rebuilt for realtime conversation
ParametersN/A (cloud)
Languages200
Self-hostable Yes
Streaming Yes
GamiWays
Phase 1 MVP — Qualité + Pipeline conversationnel complet + 100+ langues

A candidate to test for the Où est Ava ? voice layer, not the video avatar layer. Voice Direction, non-verbals and cross-turn audio context make it a direct path for how a line is delivered. Test French directions, sensitive scenes, lip-sync with the chosen renderer and end-to-end latency. For Plastic Dilemma, treat EU/India residency as an Enterprise contractual option, not an automatic sovereign deployment.

Analysis

Inworld Realtime TTS-2 is available alongside a Flash variant. It combines natural-language Voice Direction, non-verbals, cross-turn audio context, Voice Design and one voice identity across 200+ languages. TTS-2 publishes a 100 ms P90 server TTFB and Flash 20 ms; neither is an Où est Ava ? end-to-end latency. The platform also exposes STT, a Realtime API and LLM Router. French direction following, viseme compatibility and Enterprise sovereignty terms must be validated.

Strengths

  • Realtime TTS-2 disponible : direction libre et non-verbaux
  • Contexte audio inter-tours et modes Expressive/Balanced/Stable
  • TTS-2 Flash : 20 ms TTFB P90 annoncé côté serveur
  • 200+ langues et une identité vocale crosslingue annoncées
  • Clonage instantané et Voice Design selon le plan
  • Plateforme : TTS + STT + Realtime S2S + LLM Router
  • Timestamps utiles au lip-sync
  • Résidence UE/Inde et SLA/DPA à contractualiser sur Enterprise

Weaknesses

  • TTFB serveur hors réseau, LLM et vidéo
  • Direction libre et non-verbaux en français à tester scène par scène
  • Résidence UE/Inde et conditions on-premise limitées aux termes Enterprise
  • Modèles et limites du Router évoluent fréquemment

Voice Capabilities

Voice Cloning Yes

Instant voice cloning (free). Professional voice cloning (Growth/Enterprise add-on). Advanced Voice Design: create a voice from a prose description — no reference audio needed. Up to 5 custom voices (On-Demand), 100 (Creator), 1,000 (Developer), 3,000 (Growth). ElevenLabs migration tool available.

Emotion & performance directionInstruction-based direction

Available Realtime TTS-2 interprets open-ended performance directions, non-verbals and prior-turn audio context into conversational voice output.

How: Natural-language direction, non-verbal markers, cross-turn audio context and `Expressive`, `Balanced` or `Stable` modes, with voice/voice design according to plan.

Validate: Direction is a generative interpretation: measure French instruction following, reproducibility and lip-sync before an Où est Ava ? decision. Published TTFB is server-side only.

Streaming Yes

Streaming-native via WebSocket and REST. TTS-2 annonce 100 ms TTFB P90 côté serveur et Flash 20 ms P90 ; ces chiffres excluent le réseau, le LLM et la vidéo. Realtime API : WebSocket full-duplex S2S, détection de tour, tool calling et routage LLM. Les backchannels peuvent être diffusés pendant le raisonnement.

Lip-sync Data Yes

Word, character, phoneme, and viseme-level timestamps. Unity/Unreal SDKs with lipsync templates.

Pricing

Price / 1M chars
from $12.5
depending on plan
Price / minute
from $0.0125
depending on plan
Free tier
Up to 70 min TTS included on On-Demand (free)

Realtime TTS-2 : $25/1M chars On-Demand, $20 Creator, $17.50 Builder, $15 Developer et $12.50 Growth. TTS-2 Flash : $15 → $7/1M chars aux mêmes paliers. L’estimateur Inworld emploie ~1 000 caractères/min ; les crédits mensuels sont inclus dans les plans payants.

Hybrid (PAYG + Subscription)
Official pricing page
PlanSubscription/moIncludedOverage/minTop-up
On-Demand (free)

PAYG above 40 min at $25/1M chars

Free70 min$0.0250Yes
Creator

Credits rollover. Overage at $20/1M chars (TTS-2)

$25/mo$25/mo credits$0.0200Yes
GrowthRecommended

Best rate: $12.50/1M chars (TTS-2)

$1500/mo$1,500/mo credits$0.0125Yes
Top-up available: Some plans allow purchasing additional minutes (top-up) without changing plans. Useful for occasional usage spikes.

Sovereignty & Compliance

On-premise Yes

Résidence UE/Inde, SLA et DPA sont des options Enterprise. ZDR, HIPAA et BAA sont proposés en option Growth ; vérifier contractualisation, conservation des voix et éventuelles conditions on-premise séparément.

GDPR Compliant

Data residency: US, EU, India

Certifications & Compliance

SOC 2 Type II GDPR HIPAA ZDR BAA

HIPAA, ZDR (Zero Data Retention) and BAA are Growth/Enterprise options. EU or India residency and any on-premise terms must be contracted separately; none means automatic sovereign deployment.

Strategic & Business Analysis

Inworld Realtime TTS-2 + Flash — Strategic Positioning

Beyond technical specs: where does this tool sit in the ecosystem, what are the risks and strategic implications for GamiWays?

Inworld Realtime TTS-2 is now available: Voice Direction and cross-turn context make it a credible way to test how Où est Ava ? speaks. The claimed 100/20ms server TTFB and the French emotional result still require an end-to-end scene test.

Cloud + On-premise
Lock-in risk:Medium
Sovereignty fit:Medium
Open-source threat:Medium
Pricing:Stable →

A. Strategic Positioning

Target customer: Enterprise / Developer — regulated industries, gaming, real-time agents

Realtime TTS-2 available for conversation: free-form Voice Direction, Conversational Awareness, non-verbals, Voice Design and a Flash variant. Enterprise conditions govern EU/India residency and on-premise.

B. Competitive Moat

  • TTS 1.5 #1 Artificial Analysis Speech Arena (ahead of Google and ElevenLabs)
  • TTS-2 : Voice Direction — natural language delivery instructions inline (unique in market)
  • TTS-2 : Conversational Awareness — model hears prior audio turns, tone and pacing carry forward
  • TTS-2 : 200+ languages/locales announced, one voice identity and crosslingual use
  • Full platform: TTS + STT + Realtime API (WebSocket S2S) + LLM Router (220+ models)

Vulnerability: Published 100/20ms TTFB is server-side, not a full dialogue latency. French direction following, video-avatar sync and Enterprise/on-premise terms remain to be measured or contracted.

E. Strategic Questions for GamiWays

Sovereignty fit

Enterprise terms are needed for EU/India residency, SLA/DPA and any on-premise discussion. Treat ZDR, HIPAA and BAA as separately contracted features; do not classify standard access as sovereign.

Build vs. Buy

A high-priority buy/test option for Où est Ava ? voice direction, not a video-avatar purchase. Compare TTS-2 and Flash on the same French scenes; sovereignty requires an Enterprise contractual review.

Lock-in risk

Proprietary models and full-stack platform create some lock-in. Drop-in OpenAI Realtime API compatibility reduces switching cost. Competitive pricing reduces dependency risk.

Roadmap alignment

Strong alignment for the Où est Ava ? voice layer: cross-turn context may support continuity and Voice Direction can change a delivery without re-recording. Validate French instructions, non-verbals, visemes and end-to-end video timing before adoption.

This sheet's verification

Verified 6 September 2026

Inworld Realtime TTS-2 announcement and documentation, 31 Aug–6 Sep 2026. Published TTFB is server-side P90; no score or Où est Ava ? latency is inferred from it.

Update note: Realtime TTS-2 et Flash vérifiés le 6 septembre 2026 : disponibilité, TTFB P90 serveur (100/20 ms), direction libre, non-verbaux, contexte audio, 200+ langues et tarifs de $25/$15 On-Demand à $12.50/$7 Growth. Les tests français et Où est Ava ? restent à réaliser.

This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.

ELO benchmarks / indices: Artificial Analysis · Last API sync : 18 August 2026