GamiWays

Streaming Video Avatars

Shared decision cues for choosing between commercial services and open models, then understanding what actually makes a credible conversational video experience.

Strategic framing: read streaming avatars in context

What the numbers say

Advertised latency or FPS describes one part of rendering, not presence by itself. Test audio-video synchronisation, recovery after a speech turn and stability as network conditions change.

What to combine in the table

Read real-time alongside transport, conversation, body, performance direction and sovereignty. A highly realistic face may suit a demo without providing the direction, listening or scene format you require.

Stay alert to

The avatar is only one system layer: STT, TTS, LLM, WebRTC and orchestration all affect the experience. Also verify likeness rights, moderation, data residency, parallel streams, session cost and a fallback plan.

Reading method: start with the scene to sustain and the required degree of direction; then compare presence, latency and control on the same lines, before validating load, rights and fallback over a full session.

Four layers not to conflate

A living experience combines face, voice, turn-taking and data boundaries

These cues avoid mistaking a strong voice model for an avatar, or signal detection for performance direction. Each card separates what is documented from what still needs Où est Ava ? or Plastic Dilemma measurement.

Video identity and turn-taking

Tavus Phoenix-4.5 + Sparrow-2

Creates a character from image or video, declares French on the PAL and steers patience, interruption and re-engagement. Test directed French voice, perception consent/policy and full-scene latency.

Open Tavus profile →

Decoupled video rendering

HeyGen LiveAvatar LITE

Keeps avatar and video transport with HeyGen while Où est Ava ? retains STT, LLM, RAG, French voice and performance intent. Test voice switching, rights and lip-sync; no documented deterministic live gesture.

Open HeyGen profile →

How the line is delivered

Inworld Realtime TTS-2

Carries free-form direction, non-verbals, cross-turn audio context and Expressive/Balanced/Stable modes. It is a voice layer to pair with a renderer; test French directions, sensitive scenes and repeatability.

Open Inworld profile →

French, precision and residency

Gradium TTS

Targets structured terms, timestamps and stable timing. EU residency can be activated and verified on a paid plan; it is neither Swiss residency nor explicit emotion direction.

Open Gradium profile →

Two paths to compare

Same experience target, different operating constraints

Commercial services accelerate staging and live-flow testing. Open models require building the capacity but enable discussion of sovereignty, licences and infrastructure cost. Both paths use the same questions of presence, latency, direction and fallback.

Dialogue-guided generative video

For Où est Ava ?, the question is not how to rank generic clips: it is how to produce a short shot consistent with what was just said — a memory, place, clue or world shift — and insert it at the right moment. A second path explores visual environments that react live without being avatars.

A

Contextual inserts — prepare, cache or trigger

These models produce a standalone shot. They suit reactive editing; they are not a continuous video stream.

Seedance 2.5

ByteDance Seed

Insert
Access
BytePlus API announced; availability to confirm
Timing
30-second clips, extensions; asynchronous generation
Control
Up to 30 images, 10 videos and 10 audios; timestamp editing

Memory, location or clue shot triggered between two exchanges

Official source ↗

Gemini Omni Flash + Veo 3.1

Google

Insert
Access
Gemini API
Timing
Generative job; not a frame-by-frame stream
Control
Text, image, audio and video references; multi-turn editing; extension / last-frame control with Veo

Versioned insert factory from dialogue context

Official source ↗

Grok Imagine API

xAI

Insert
Access
API and partner platforms
Timing
Rapid iteration, to be measured on the Où est Ava ? scenario
Control
Text / image, object, motion and scene editing

Preview or modify an insert while retaining the source shot

Official source ↗
B

Interactive visual streams — non-avatar

These world models generate images as an action, navigation or event unfolds. They open a path to reactive settings or worlds, with their own costs, rights and action limits.

Runway GWM-Worlds

Runway · Early access / to qualify

Real-time frame-by-frame generation

Camera pose, actions and audio; explorable environment

Explored scene or reactive setting, separate from the avatar

Official source ↗

Google Genie 3

Google DeepMind · Research prototype / limited Project Genie access

720p at 20–24 fps; a few minutes of interaction

Navigation and prompt-controlled world events

Research path for a reactive world, not a production building block

Official source ↗

World Labs RTFM

World Labs · Research preview

Interactive frame rate targeted on one H100

Interactive frames and spatial persistence from images

Spatial path to cost before any service promise

Official source ↗

Decart Lucy / Oasis

Decart · Product access to qualify

Lucy: announced live transformation at 30 fps

Live stream transformation or interactive world depending on product

Test a consented live transformation, not a generated narrative insert

Official source ↗

Business & Market

Two recent reference points to read separately: market definitions are not additive.

Digital humans · global

$7.96B · 2026→ $26.04B · 203126.76% CAGR

Interactive avatars represented 58.63% of the segment in 2025.

Mordor Intelligence ↗

AI avatars · United States

$287.0M · 2026→ $1.84B · 203330.4% CAGR

Avatar-based customer service represented 32.9% of this market in 2025.

Grand View Research ↗

Digital humans · global

USD billions · published checkpoints, not annual interpolation

2025$6.28B2026$7.96B2031$26.04B● observed / estimated○ forecast

AI avatars · United States

USD millions · published checkpoints, not annual interpolation

2025$218.7M2026$287M2033$1.84B● observed / estimated○ forecast

Structure indicators · digital humans

Mordor Intelligence ↗
Interactive share58.63%

Digital humans · 2025

Cloud deployment71.12%

Digital humans · 2025

GenAI systems46.41%

Digital humans · 2025

Reading for GamiWays: segment growth supports exploration, but does not establish the feasibility of Où est Ava ? or Plastic Dilemma. The decision criterion remains evidence of a credible, controllable and economically sustainable field experience.

GamiWays strategic positioning

GamiWays combines AI conversation, video avatar, sequencing and narrative or learning control. Value does not come from an isolated video model: it depends on integration, the chosen sovereignty posture and measurements made in real experiences.