Voice-to-Voice Pipeline
Interactive diagram of the complete voice pipeline for GamiWays Phase 1 MVP. Select components for each block to visualize cumulative latency and estimated cost. Compare the Cascade approach (ASR → LLM → TTS) with end-to-end Voice-to-Voice.
Estimated Latency & Cost
GamiWays Phase 1 target: <2s end-to-end (voice pipeline only, excluding avatar generation). Avatar generation adds 80–300ms (BeyondPresence) or 3–8s (HeyGen).
Converts audio stream to text. Streaming ASR sends partial transcripts to reduce LLM Time-to-First-Token. Critical latency block.
Industry reference for real-time ASR. Supports 30+ languages. Used by most voice agent frameworks.
Best sovereign option. Use faster-whisper with streaming VAD for near-real-time. Quantized small.en: ~200ms on CPU.
Débit batch publié : RTFx 3332,74. WER moyen : 6,34% sur HF Open ASR Leaderboard ; instantané du rapport : 6,32% vs 7,44% pour Whisper large-v3 sur le même protocole. NeMo propose chunks 2s + contexte droit 2s pour l’intégration, mais aucune TTFA P50/P95 n’est publiée : la latence est donc à mesurer. Ne pas confondre TDT v3 et le microservice Parakeet CTC/Riva distinct.
Good alternative with EU data residency option. Better for multi-speaker scenarios.
Gradium STT with semantic VAD enables meaning-based turn detection (not just silence). Word-level timestamps for avatar lip-sync. Native LiveKit/Pipecat integration. On-premise Enterprise with zero data retention. WER to validate independently via Pipecat benchmark.
Key advantage: when combined with Inworld TTS and LLM Router, eliminates inter-component network hops. Semantic VAD reduces hallucination triggers. Use for Inworld Single-Provider stack.
ASR → LLM → TTS — modular, controllable, production-ready
Recommended for MVP Phase 1. Best stack: Deepgram Nova-3 + GPT-4o streaming + Cartesia Sonic 3. Sovereign alternative: Whisper.cpp + Llama 3.1 8B + Kokoro 82M.
Direct audio-in → audio-out — lowest latency, natural prosody
Recommended for R&D Axis 1 exploration (H2 2026). Not suitable for Phase 1 MVP due to lack of voice cloning. Monitor Voxtral TTS (Mistral) and Ultravox v0.5.
Recommended Stacks
Click 'Apply' to load components into the configurator.
MVP Cloud Stack
Fastest path to working prototype
Phase 1 prototype path. Deepgram (75ms) + GPT-4o streaming (350ms) + Sonic 3.6 (under 90ms, Cartesia figure) + WebRTC (30ms) ≈ 555ms as a design sum, not an Où est Ava ? measurement. Ink-2 is not the French STT in this stack.
Sovereign Stack
Full Swiss sovereignty — Exoscale/OVH deployment
Full sovereignty for Swiss/EU institutional partners. Whisper.cpp (200ms) + Llama 3.1 8B (150ms) + Chatterbox (150ms) + Mem0 (20ms) + WebRTC (30ms) = ~710ms best-case. Voice cloning via Chatterbox. Deployable on Exoscale Geneva.
Inworld Conversation Layer
STT + Router + TTS-2 can reduce hand-offs; validate the complete path
Inworld STT, Router and Realtime TTS-2 may reduce application-level hand-offs when they are used together; this does not remove network, memory, external LLM or video-rendering time. TTS-2 adds free-form Voice Direction, non-verbals and cross-turn audio context; Flash publishes a 20ms server TTFB. The 470ms best-case is only a scenario calculation, not an Où est Ava ? result. EU/India residency and on-premise require Enterprise terms. Test French directions, avatar lip-sync and full end-to-end behaviour.
Hybrid Stack
Best quality/sovereignty balance for production
EU-sovereign LLM (Mistral) + sovereign TTS (Kokoro) + best ASR (Deepgram) + Mem0 long-term memory. Deepgram (75ms) + Mistral Nemo (200ms) + Kokoro (60ms) + Mem0 (20ms) + WebRTC (30ms) = ~555ms best-case. No voice cloning — add Chatterbox for persona.
Key decision: voice cloning
Voice cloning is critical for GamiWays persona. Options: Cartesia (cloud, 40ms), Chatterbox (local, 150ms, MIT), ElevenLabs (cloud, 75ms, $75/1M). Sovereign stack requires Chatterbox + Kokoro combination.
Compare all TTS solutions
14 TTS/V2V solutions with comparative scores on 7 axes.
TTS State of the Art →