GamiWays

Avatar Behavior & Live Video Experience

Structuring the challenges for a credible streaming video experience — a living stream, not just a talking face.

The central challenge: a live video stream, not TTS with a face

GamiWays' goal is not to produce an audio avatar with synchronized facial animation — it is to deliver a continuous, reactive video experience where the user perceives a real presence. This requires a Stream B able to translate a performance direction attached to the line — intention, intensity, pace, silence, closeness — into coherent voice and presence. TTS remains a pipeline component, but the final output is video and the how matters as much as the text.

Stream A: Video analysis — offline, non-criticalStream B: Avatar video stream — main real-time challengeKPI: TTFA < 500ms · Fluidity 25fps+ · Audio/video sync < 100ms
STREAM A — Source Video AnalysisOffline processing · Not a major R&D challengeVideo Archives(source material)✓ Offline · SimpleFrameExtractionJPEG · 1fpsSemanticAnalysisCLIP · BLIP2Tags · EmbeddingsVideoDescriptor DBVector DBSemantic searchDynamic VideoPlaylistUpdated in real-timebased on conversationIllustrative VideosPlay alongside avatar(secondary stream)Avatar speaks ✓Informative / Interview VideosAvatar pausesvideo delivers spoken contentAvatar pauses ⏸STREAM B — Avatar ConstructionOffline training + real-time inference · Main R&D challengeTrainingVideoSingle · Never playedin the experienceBehavioralFingerprintMicro-expressionsGestures · RhythmProsodyAvatarModel⚠ R&D Challenge (Axis 2b)Diffusion distillationIntelligent cache<500ms targetReal-timeAvatarSpeaks · Pauses duringinformative videosOUTPUT — Dual-Stream ExperienceMulti-stream sync · <100ms · Memoways internal expertiseMain StreamReal-time avatarWebRTC · H.264 · <100msPauses during informative videosSecondary StreamDynamic video playlistIllustrative: insert alongside avatarInformative: full-screen, avatar pausesSmart orchestration: avatar yields to informative videosIllustrative videos play as inserts without interrupting avatarResearch Axes: Axis 2b (Avatar <500ms) · Axis 2a (Expressive TTS) · Axis 1 (Conversational Memory)Source video analysis (Stream A) is NOT a research challenge — standard offline processing
↔ Zoom
Click to expand

The four dimensions of a living video experience

For a video avatar to be perceived as alive and credible, four dimensions must be addressed simultaneously — not sequentially.

Behavioral extraction from archives

Extract individual behavioral patterns from existing videos — without new capture sessions. Identify micro-expression repertoire, gestural vocabulary, gesture-speech temporal relationships, and postural habits.

Key question:

Can we automatically extract an individual's gestural vocabulary from uncontrolled footage?

Coherent and coordinated body language

Go beyond lip-sync. Generate coordinated body behavior: synchronized with speech content and emotional tone, culturally appropriate, consistent with the defined personality. Most current systems focus on the face only — the body is absent or from a template library.

Key question:

How to ensure body language, facial expressions, and prosody tell the same emotional story simultaneously?

Live and credible video streaming

The target experience is not an audio avatar with an animated face — it is a continuous, reactive video stream that conveys the impression of a real presence. This requires real-time rendering of face and body, synchronized with generated speech, without perceptible latency or visual artifacts. TTFA (Time to First Audio/Frame) and stream fluidity are the primary KPIs.

Key question:

Can a smooth avatar video stream be maintained under 500ms total latency on a standard network connection?

Real-time performance under latency constraints

Approaches: pre-rendered base + real-time lip-sync, model distillation, intelligent cache, graceful degradation. The goal is an acceptable personalized video avatar at under 500ms on accessible hardware. Simli Trinity-1 (≤202F300ms TTFA) and bitHuman Essence (CPU, no GPU) represent the two extremes of the latency/cost spectrum.

Key question:

Can quantization, pruning, or distillation make real-time video generation feasible on accessible GPU (A10G)?

The emotional layer: what makes the experience hold up over time

A conversational video avatar is not just a smooth video stream. Behavioral fidelity requires an explicit emotional design layer: defining, encoding, and activating a repertoire of states consistent with the character's personality, history, and interaction context.

Emotional repertoire

Define a set of discrete and continuous emotional states per character. Each state encodes: facial expression, vocal prosody, cadence, posture, micro-behaviors. The repertoire is the foundation of all emotional coherence over time.

Transitions and temporal coherence

Transitions between emotional states must be smooth, personality-consistent, and not create perceptible breaks. The challenge is avoiding the 'emotional uncanny valley' — when the avatar shifts state abruptly or inconsistently with its history.

Real-time contextual activation

Emotional state draws on content, history and incoming cues (tone, rhythm, silence). Perceiving cues must be separated from responding: Hume EVI-3 documents prosodic perception, while Inworld Realtime TTS-2 can carry vocal direction, non-verbals and prior-turn audio context. Neither replaces an end-to-end Où est Ava ? measurement.

From text to performance: what can actually be controlled

Direction should travel with each line in structured form: dramatic intent, energy, pace, pauses, authorised non-verbal cues and visual state. Providers do not receive this information in the same way: some expose explicit API parameters, others interpret tags or a prompt, and others can only be animated by the audio supplied as input.

Explicit control

For turn-by-turn direction: Anam (director notes and cues), Simli (facial state) and bitHuman (prepared gestures). On voice, Cartesia, Hume and Flux expose more direct controls.

Interpreted direction

Tags and prompts from ElevenLabs, Inworld, xAI, Tavus or LemonSlice can steer performance, but must be evaluated line by line: they are not deterministic visual commands.

Audio-driven: useful fallback

For Hedra, Avatario or MuseTalk, emotion must come from expressive TTS or pre-produced audio. This is a fallback option, not real-time staging control.

Storygami

Streaming video solutions — Conversational cinema

Storygami requires a cinematic video stream: latency < 500ms imperative, native full-duplex, creator-controllable emotional expressivity. The avatar embodies a character with a biography — every fluidity break destroys immersion.

Lowest market latency (< 300ms TTFA)
Native full-duplex — interruptions handled
Real-time WebSocket, continuous streaming
Creator API with partial emotional control
Facial fidelity lower than HeyGen Avatar V
Limited voice cloning (pre-defined voices)
Price: $0.10/min (Pro) — costly at scale
Parallel streams limited by plan

Best choice for Storygami: cinematic latency, full-duplex, continuous streaming. Tradeoff on facial fidelity.

Ultra-low latency (< 200ms) — CPU only, no GPU
Continuous streaming, no chunk generation
On-premise deployment possible (sovereignty)
Very low infrastructure cost (CPU cloud)
Limited visual fidelity vs GPU solutions
Less rich emotional expressivity
Fewer available voices
Less mature API documentation

Economic alternative for Storygami when latency takes priority over fidelity. Ideal for sovereign deployments.

Premium video quality — full body possible
Native emotional intelligence (Genesis 2.0)
Full-duplex with interruption handling
Avatar cloning from existing video
Higher latency than Simli (< 500ms vs < 300ms)
Premium price — not publicly documented
Parallel streams limited in standard plan
Less creator control over emotions

High-fidelity option for Storygami when visual quality takes priority. Acceptable latency for less reactive scenes.

Emotion API + Action API — explicit creator control
Multi-style: 1 actor covers multiple characters
Self-Managed Pipeline available (A100/H100)
Continuous streaming with audio/video sync
Intermediate latency (< 400ms)
High cloud price ($0.21/min)
Requires powerful GPU for self-hosting
Less accessible documentation

Best creator control for Storygami. Only platform with explicit Emotion API. Cost to monitor.

To measure in Où est Ava ?
Video CVI: Phoenix-4.5 creates a face from image or video, while Raven-1 reads audio-visual cues according to policy
Sparrow-2 considers pauses, prosody, interruptions, backchannels, noise and background speech
Patience, interruptibility and idle engagement are configurable in conversational flow
Custom replica, French PAL and Echo mode make it possible to test directed external French voice
Tavus does not publish Où est Ava ? end-to-end dialogue latency: it must be measured
These settings steer turn-taking, not a determined gesture, gaze or expression
US hosting; consent and EU policy must be verified if Raven-1 is used
First paid plan: $22/60 min and only one concurrent session

Sparrow-2 criteria · turn-taking

A test candidate for scenes where knowing when to wait matters as much as replying: silence after a heavy event, a backchannel or overlapping speech. Evaluate it with Où est Ava ?’s French voice, network and performance direction before making any fluidity claim.

Edugami

Streaming video solutions — Pedagogical agent

Edugami prioritizes benevolent and adaptive presence: the avatar detects learner frustration or confusion. Latency < 800ms is acceptable, but data sovereignty (GDPR) and cost per session are strong constraints.

< 500ms
Native emotional intelligence — frustration/confusion detection
Designed for pedagogical conversational agents
Accessible price: Starter $12/month (30 min included)
Integrated body language, not just face
Emotional control not configurable by creator
Limited parallel streams (Free: 1, Starter: 3)
Limited multilingual (EN/FR mainly)
Latency < 500ms — acceptable but not cinematic

Best choice for Edugami: native pedagogical agent, accessible price, adaptive emotions. Ideal for Dilemme Plastique.

Raven-1: incoming perception subject to learner policy and consent
Phoenix-4.5: character created from image or video, then to validate for a stable mediator
Sparrow-2: pauses, interruptions and noise considered in conversational flow
Personalized learner responses with participant attribution
Total dialogue latency is not published: measure it with the educational stack
First paid plan: $22/60 min; Builder overage $0.35/min
Concurrent sessions: Starter 1, Builder 3, Growth 10
Turn-taking settings are not deterministic gesture direction

A test candidate for attentive pedagogical dialogue, without promising target latency or authorised emotion perception. First verify pauses, interruptions, consent/policy, cohort cost and mediator stability before classroom use.

Avatar cloning from single photo — very accessible
Well-documented API, easy integration
Accessible price in Lite plan ($5.9/month)
Multilingual (40+ languages)
Latency 3–6s — incompatible with real-time conversation
No full-duplex, no continuous streaming
Limited emotional expressivity
Parallel streams not publicly documented

Suitable only for Edugami in async mode (pre-generated educational videos). Not viable for real-time.

Audio fallback solutions — if video is unavailable or too costly

In cases where video streaming is not viable (network constraints, budget, strict GDPR), these audio solutions maintain credible emotional presence through voice alone. Ranked by relevance for Edugami.

~$0.043/min (XS)

A French option for terminology precision, timestamps and stable timing: Gradium reports 216ms P50 TTFA and 81% on structured entities including French. EU residency must be enabled on a paid plan and verified; it does not by itself make deployment sovereign. No explicit emotion control is documented.

$0.003/min

Native sound tags ([laughter], [sigh], [hesitation]) simulate emotional presence without video. Prompt-controllable prosodic expressivity. Voice cloning from 1 min of audio. High-quality fallback if video is unavailable.

$0.005/min

Sonic 3.6 (27 August 2026): 44 languages, under 90 ms as claimed by Cartesia. Ink-2 stays English-only, so it is not the French STT in the pair. The portal ELO is still the 18 August snapshot, not a new rank.

Key differentiator — state of the art August 2026

The state of the art is more nuanced than a simple “expressive / not expressive” label. Anam currently exposes the most detailed real-time direction (director notes, intensity and synchronised cues); Simli exposes discrete facial states; bitHuman triggers prepared gestures. No evaluated service documents, in streaming, a deterministic per-turn command for prosody, gaze, gesture, framing and expression at once. The Où est Ava ? decision is therefore a test of a voice + avatar + orchestrator pipeline, not the choice of a magic avatar.

Comparison — Avatar video streaming platforms

Click a header to sort

PlatformTTFALip-syncMax streamsPriceSovereigntyTarget use
< 200ms✓ CPU50 (Business)$0.01/minUS — On-prem possibleStorygami ✓ / Edugami ✓
< 300ms✓ Native10 (Pro)$0.10/minUS — AWSStorygami ✓✓
< 400ms✓ Multi-style3 (Starter)$0.21/minUS — Self-hosted A100Storygami ✓✓
< 500ms✓ + Body3 (Starter)$0.12/minEU — GDPREdugami ✓✓
< 500ms✓ Full body10 (Starter)On requestEU — On-prem EnterpriseStorygami ✓✓
< 500ms✓ Open-sourceGPU-dependentGPU cloud ~$0.05/minSelf-hosted — sovereignStorygami/Edugami self-hosted
2–5s✓ High fidelity20 (Essential)$0.05/secUS — AWSAsync only
3–6s✓ Photo → videoNot documented$5.9/moisUS — AWSEdugami async
CVI: to measure✓ High fidelity1 Starter · 3 Builder · 10 Growth$0.367/min (Starter)US — AWSOù est Ava ?/Edugami: test needed

Updated: August 2026 — New = added · Upd = updated · TTFA colored: < 300ms · < 500ms · > 1s