Shared decision cues for choosing between commercial services and open models, then understanding what actually makes a credible conversational video experience.
🎯
Strategic framing: read streaming avatars in context
What the numbers say
Advertised latency or FPS describes one part of rendering, not presence by itself. Test audio-video synchronisation, recovery after a speech turn and stability as network conditions change.
What to combine in the table
Read real-time alongside transport, conversation, body, performance direction and sovereignty. A highly realistic face may suit a demo without providing the direction, listening or scene format you require.
Stay alert to
The avatar is only one system layer: STT, TTS, LLM, WebRTC and orchestration all affect the experience. Also verify likeness rights, moderation, data residency, parallel streams, session cost and a fallback plan.
Reading method: start with the scene to sustain and the required degree of direction; then compare presence, latency and control on the same lines, before validating load, rights and fallback over a full session.
Four layers not to conflate
A living experience combines face, voice, turn-taking and data boundaries
These cues avoid mistaking a strong voice model for an avatar, or signal detection for performance direction. Each card separates what is documented from what still needs Où est Ava ? or Plastic Dilemma measurement.
Video identity and turn-taking
Tavus Phoenix-4.5 + Sparrow-2
Creates a character from image or video, declares French on the PAL and steers patience, interruption and re-engagement. Test directed French voice, perception consent/policy and full-scene latency.
Keeps avatar and video transport with HeyGen while Où est Ava ? retains STT, LLM, RAG, French voice and performance intent. Test voice switching, rights and lip-sync; no documented deterministic live gesture.
Carries free-form direction, non-verbals, cross-turn audio context and Expressive/Balanced/Stable modes. It is a voice layer to pair with a renderer; test French directions, sensitive scenes and repeatability.
Targets structured terms, timestamps and stable timing. EU residency can be activated and verified on a paid plan; it is neither Swiss residency nor explicit emotion direction.
Same experience target, different operating constraints
Commercial services accelerate staging and live-flow testing. Open models require building the capacity but enable discussion of sovereignty, licences and infrastructure cost. Both paths use the same questions of presence, latency, direction and fallback.
For Où est Ava ?, the question is not how to rank generic clips: it is how to produce a short shot consistent with what was just said — a memory, place, clue or world shift — and insert it at the right moment. A second path explores visual environments that react live without being avatars.
A
Contextual inserts — prepare, cache or trigger
These models produce a standalone shot. They suit reactive editing; they are not a continuous video stream.
These world models generate images as an action, navigation or event unfolds. They open a path to reactive settings or worlds, with their own costs, rights and action limits.
Runway GWM-Worlds
Runway · Early access / to qualify
Real-time frame-by-frame generation
Camera pose, actions and audio; explorable environment
Explored scene or reactive setting, separate from the avatar
Reading for GamiWays: segment growth supports exploration, but does not establish the feasibility of Où est Ava ? or Plastic Dilemma. The decision criterion remains evidence of a credible, controllable and economically sustainable field experience.
GamiWays strategic positioning
GamiWays combines AI conversation, video avatar, sequencing and narrative or learning control. Value does not come from an isolated video model: it depends on integration, the chosen sovereignty posture and measurements made in real experiences.