GamiWays

Voice-to-Voice & Real Time

Compare global voice architectures: modular cascade, integrated platforms and speech-to-speech models. Detailed recognition and synthesis comparisons remain on their dedicated pages.

Strategic framing: read Voice-to-Voice in context

The metric that counts

Measure end-of-speech → first audio, then review median, P95 and jitter. A vendor latency or isolated first audio does not describe felt silences, recoveries and variation.

The infrastructure layer

WebRTC, turn detection, barge-in, audio routing, reconnection and session memory determine continuity. An S2S model replaces neither this control layer nor its observability.

Stay alert to

Test concurrent streams, backpressure, session cost, data residency and fallback to a controlled cascade. The goal is a stable conversation, not only a fast demo.

Reading method: set the required control level first, then compare complete latency on the same network, and validate interruptions, load and fallback before selecting an architecture.

Two architectures to compare

01

Cascade Pipeline

ASR → LLM → TTS — modular, controllable, production-ready

505–1150ms

What it brings

  • Full control over each component
  • Swap any block independently
  • Voice cloning via TTS layer

What to trade off

  • Cumulative latency (3 sequential steps)
  • Prosody lost between LLM and TTS
  • Requires orchestration layer
02

End-to-End Voice-to-Voice

Direct audio-in → audio-out — lowest latency, natural prosody

150–350ms

What it brings

  • Lowest possible latency (150–350ms)
  • Natural prosody preserved end-to-end
  • No text bottleneck

What to trade off

  • No voice cloning (critical for GamiWays)
  • Limited control over content/persona
  • Models still maturing (2025–2026)

Values are design references: each experience must measure real latency from end of speech to first audio received.

Voice-to-Voice paths to evaluate

References from existing voice sheets. Open a row to review advantages, limits and test conditions.

Hume EVI 3

Speech-to-speech

< 300 ms claimed

Conversational audio loop with a particular strength in emotional intent. Test it when delivery matters more than modular composition.

Control
High — natural-language emotion
Sovereignty
Low — US cloud
Moshi (Kyutai)

Full-duplex · R&D

200–400 ms reported

A research reference for speaking and listening simultaneously, especially useful for interruptions, backchannels and natural turn-taking.

Control
Limited — full-duplex research
Sovereignty
High — open weights
Inworld Realtime API

Integrated platform

≈ 490 ms best case*

Unifies STT, LLM routing and streaming TTS-2. It is not a pure audio-to-audio model, but an integrated stack that reduces network hops between components.

Control
High — Voice Direction + tools
Sovereignty
Partial — Enterprise on-premise
≈ 0.70 s first audio

A WebSocket speech-to-speech stack that combines transcription, reasoning and audio response. Useful for measuring the benefit of an integrated path against a modular cascade.

Control
Partial — qualify TTS tags in S2S
Sovereignty
Low — US cloud
Ultravox v0.5

Open-weights speech-to-speech

0.864 s referenced median

An open-weights path to test for sovereignty and reduced serialization, without attributing to it the control guarantees of a cascade.

Control
Limited — less decomposable output
Sovereignty
High — self-hostable
OpenAI gpt-realtime-2.1

Speech-to-speech

1.536 s referenced median

gpt-realtime-2.1 continues the speech-to-speech path with, according to OpenAI, better interruptions, noise and alphanumerics. Audio tokens stay at $32 / $64 per million. The mini variant is cheaper. TurnBench below still measures the documented VADs, not Live-1.

Control
Partial — tools + reasoning
Sovereignty
Partial — EU residency
Milliseconds not published

Full-duplex voice front end: it listens and speaks at once, and delegates tools and reasoning to a backend model. $0.05/min for that layer, plus the model behind it. The announced turn-taking gain versus gpt-realtime-2.1 is an OpenAI figure. Useful as an architecture comparison, not as a replacement for the French cascade.

Control
Tone by prompt; reasoning is delegated
Sovereignty
Low — cloud, residency still to retest

The decisive link: real-time infrastructure

A Voice-to-Voice model does not replace WebRTC, interruption handling or observability. Track end-of-speech → first audio, barge-in, jitter, concurrent streams and controlled fallback when a service slows down.

Compare real-time infrastructure responses

Decide by scenario, then measure

For Où est Ava ?, cascade remains the reference when persona, wording and voice require control. S2S paths are comparative experiments for fluidity and emotion. Decide from end-to-end tests, not from one vendor latency claim.

Configure a stack and estimate its latency