GamiWays
PHASE 1 MVPReference architecture

Voice-to-Voice Pipeline

Interactive diagram of the complete voice pipeline for GamiWays Phase 1 MVP. Select components for each block to visualize cumulative latency and estimated cost. Compare the Cascade approach (ASR → LLM → TTS) with end-to-end Voice-to-Voice.

Estimated Latency & Cost

Best case
575ms
Typical
1250ms
Cost/min
$0.082
Cost/hour
$4.94
✓ Cible <2s
Best case
Typical
2s target
User: 10ms · free
ASR: 75ms · $0.004/min
Memory: 20ms · $0.004/min
LLM: 350ms · $0.060/min
TTS: 90ms · $0.014/min
Transport: 30ms · free
Cost breakdown
■ ASR: 5%■ Memory: 5%■ LLM: 73%■ TTS: 17%

GamiWays Phase 1 target: <2s end-to-end (voice pipeline only, excluding avatar generation). Avatar generation adds 80–300ms (BeyondPresence) or 3–8s (HeyGen).

Pipeline flow

Converts audio stream to text. Streaming ASR sends partial transcripts to reduce LLM Time-to-First-Token. Critical latency block.

Deepgram Nova-3
RecommandéCloud
75–200ms
$0.0043/min streaming

Industry reference for real-time ASR. Supports 30+ languages. Used by most voice agent frameworks.

Whisper.cpp (local)
AlternatifLocal
200–500ms
Free (MIT)
Souverain

Best sovereign option. Use faster-whisper with streaming VAD for near-real-time. Quantized small.en: ~200ms on CPU.

NVIDIA Parakeet TDT 0.6B v3
AlternatifLocal
to measure
Open weights (CC-BY-4.0) — GPU / CapEx à mesurer
Souverain

Débit batch publié : RTFx 3332,74. WER moyen : 6,34% sur HF Open ASR Leaderboard ; instantané du rapport : 6,32% vs 7,44% pour Whisper large-v3 sur le même protocole. NeMo propose chunks 2s + contexte droit 2s pour l’intégration, mais aucune TTFA P50/P95 n’est publiée : la latence est donc à mesurer. Ne pas confondre TDT v3 et le microservice Parakeet CTC/Riva distinct.

AssemblyAI Universal-2
AlternatifCloud
100–250ms
$0.0065/min streaming

Good alternative with EU data residency option. Better for multi-speaker scenarios.

Gradium STT
AlternatifCloud
100–200ms
$0.009/min (45k crédits/mois gratuits)

Gradium STT with semantic VAD enables meaning-based turn detection (not just silence). Word-level timestamps for avatar lip-sync. Native LiveKit/Pipecat integration. On-premise Enterprise with zero data retention. WER to validate independently via Pipecat benchmark.

Inworld STT
AlternatifCloud
80–180ms
Included in Inworld platform (TTS + STT + Realtime API bundle)
Souverain

Key advantage: when combined with Inworld TTS and LLM Router, eliminates inter-component network hops. Semantic VAD reduces hallucination triggers. Use for Inworld Single-Provider stack.

Cascade Pipeline

ASR → LLM → TTS — modular, controllable, production-ready

Best
505ms
Typical
1150ms

Recommended for MVP Phase 1. Best stack: Deepgram Nova-3 + GPT-4o streaming + Cartesia Sonic 3. Sovereign alternative: Whisper.cpp + Llama 3.1 8B + Kokoro 82M.

End-to-End Voice-to-Voice

Direct audio-in → audio-out — lowest latency, natural prosody

Best
150ms
Typical
350ms

Recommended for R&D Axis 1 exploration (H2 2026). Not suitable for Phase 1 MVP due to lack of voice cloning. Monitor Voxtral TTS (Mistral) and Ultravox v0.5.

Recommended Stacks

Click 'Apply' to load components into the configurator.

MVP RECOMMENDED

MVP Cloud Stack

Fastest path to working prototype

Best latency
555ms
Typical
1140ms
Cost/min
$0.085
US Cloud — acceptable for prototype

Phase 1 prototype path. Deepgram (75ms) + GPT-4o streaming (350ms) + Sonic 3.6 (under 90ms, Cartesia figure) + WebRTC (30ms) ≈ 555ms as a design sum, not an Où est Ava ? measurement. Ink-2 is not the French STT in this stack.

Sovereign Stack

Full Swiss sovereignty — Exoscale/OVH deployment

Best latency
710ms
Typical
1620ms
Cost/min
$0.015
Full sovereignty — Exoscale/OVH deployable

Full sovereignty for Swiss/EU institutional partners. Whisper.cpp (200ms) + Llama 3.1 8B (150ms) + Chatterbox (150ms) + Mem0 (20ms) + WebRTC (30ms) = ~710ms best-case. Voice cloning via Chatterbox. Deployable on Exoscale Geneva.

MVP RECOMMENDED

Inworld Conversation Layer

STT + Router + TTS-2 can reduce hand-offs; validate the complete path

Best latency
470ms
Typical
1050ms
Cost/min
$0.060
US Cloud — acceptable for prototype

Inworld STT, Router and Realtime TTS-2 may reduce application-level hand-offs when they are used together; this does not remove network, memory, external LLM or video-rendering time. TTS-2 adds free-form Voice Direction, non-verbals and cross-turn audio context; Flash publishes a 20ms server TTFB. The 470ms best-case is only a scenario calculation, not an Où est Ava ? result. EU/India residency and on-premise require Enterprise terms. Test French directions, avatar lip-sync and full end-to-end behaviour.

Hybrid Stack

Best quality/sovereignty balance for production

Best latency
555ms
Typical
1230ms
Cost/min
$0.030
US Cloud — acceptable for prototype

EU-sovereign LLM (Mistral) + sovereign TTS (Kokoro) + best ASR (Deepgram) + Mem0 long-term memory. Deepgram (75ms) + Mistral Nemo (200ms) + Kokoro (60ms) + Mem0 (20ms) + WebRTC (30ms) = ~555ms best-case. No voice cloning — add Chatterbox for persona.

Key decision: voice cloning

Voice cloning is critical for GamiWays persona. Options: Cartesia (cloud, 40ms), Chatterbox (local, 150ms, MIT), ElevenLabs (cloud, 75ms, $75/1M). Sovereign stack requires Chatterbox + Kokoro combination.

Compare all TTS solutions

14 TTS/V2V solutions with comparative scores on 7 axes.

TTS State of the Art →