GamiWays
← GamiWays Project›State of the art
Strategic watch2024–2026

State of the Art & Research Perspectives

Research that informs GamiWays’ next decisions: reliable conversational memory, vocal acting direction, credible avatar presence, and controllable orchestration.

Key insight: each reference is linked to an Où est Ava ?, Storygami, Edugami, or Core protocol to move from an academic result to a measurable hypothesis.

01

Research decisions

Section

Direct relevance

What research unlocks right now

Four Phase B decisions to inform before multiplying integrations or demonstrators.

01

Measure useful memory

Separate what must be retained, summarized, contradicted, or forgotten across sessions — rather than merely extending context.

02

Control how, not only what

Evaluate intent, pace, silence, and prosody scene by scene so characters do not merely recite a response.

03

Test video presence

Measure speaking, listening, non-verbal response, and identity stability separately instead of scoring the avatar globally.

04

Arbitrate cost, sovereignty, and trust

Compare architectures on perceived latency, infrastructure requirements, data rights, and the ability to fall back to audio.

02

Strategic synthesis

Section

Decision matrix

Prioritize what must be validated, decided, or secured

Choose a dimension, then narrow the list by impact or context. Each issue opens a research reading, a priority action, and the next paths to pursue.

How to use it: select an issue to reveal the recommended action.

Filter by impact

Filter by context

6 issues shown

03

Corpus & protocols

Section
Browsable corpus

Find research by the decision it can inform

Filter references by the issue at hand instead of scanning a linear bibliography. Each card connects a result to a concrete hypothesis, test, or trade-off for the Core, Storygami, or Edugami.

45 references shown

Keywords: audio-avatar LLM, expressive facial movements, LoRA architecture

Video avatarPhase B
▶ Architecture: 8B parameters + 0.16B LoRA

A²-LLM is an end-to-end audio-avatar LLM that generates emotionally rich facial movements beyond lip-sync, by directly coupling text generation and facial synthesis in a single model. The 8B + 0.16B LoRA architecture enables fine adaptation to different expressive styles without retraining the entire model.

Keywords: streaming diffusion, Local-Future Sliding-Window Denoising, long-form real-time

Video avatarReal timePhase B
▶ Innovation: single image + streaming audio → real-time long-form video in 1 diffusion step

AvatarForcing is the most directly applicable academic paper to R&D Axis 1 (latency reduction). Its 'Local-Future Sliding-Window Denoising' technique enables talking avatar video generation in 1-step streaming diffusion — drastically reducing denoising iterations (from 20–50 steps to 1 step), dividing generation time by 60–90%.

Architecture: MoE 276B/12B active, native dMel audio tokenization, dual Fast/Slow model

Phase B
▶ Status: Waitlist preview, May 2026 — no public API

TML Interaction-Small is a MoE model with 276B total parameters but only 12B active per inference. It integrates native dMel audio tokenization and a dual architecture: the Fast Model (200ms chunks) handles real-time conversation while the Slow Model (asynchronous) handles tool calls, RAG, and deep reasoning. Visual proactivity and absolute temporal awareness. Known limitation: coherence degrading after ~30–45 minutes.

Keywords: full-duplex, asynchronous RAG, Moshi, enriched memory

MemoryReal timePhase B
▶ Target latency: 160–260ms + contextual enrichment

MoshiRAG is an extension of Moshi (Kyutai) integrating asynchronous RAG and enriched memory in a full-duplex model. The target architecture is exactly GamiWays' Phase C: full-duplex + RAG/retrieval/reasoning in parallel, without blocking the audio stream. The GamiWays Core Engine is architecturally this orchestration layer.

Keywords: duplex cascade, conversational pipeline, reduced latency, streaming

Phase BReal time

DuplexCascade proposes a duplex cascade conversational pipeline architecture, reducing perceived latency by starting each step (TTS, avatar generation) from the first tokens of the previous step, without waiting for completion. This is the academic formalization of the 'streaming pipeline' approach GamiWays targets for Phase A+ (estimated gain: −40% total latency).

Keywords: real-time / sync / emotion trilemma, 3DMM, single-pass GAN, 55+ FPS

Video avatarReal time
▶ Innovation: first framework to simultaneously resolve all 3 constraints — 55+ FPS, lip sync, label-free emotional fidelity

RETA is the first framework to simultaneously resolve the fundamental trilemma: real-time performance, lip synchronization accuracy, and emotional fidelity. It disentangles the audio signal into two representations: a 3DMM geometry for lip synchronization, and a dynamic emotional embedding learned without labels via cross-modal distillation. A single-pass GAN generator integrates these representations hierarchically. 55+ FPS with new SOTA on all metrics.

Keywords: Local-Future Sliding-Window, 1.3B parameters, 34 ms/frame, 2-step distillation

Video avatarReal time
▶ Innovation: 34 ms/frame (≈29 FPS), 1.3B parameters — lightest and most actionable architecture for production

AvatarForcing resolves the exposure bias of autoregressive models via a sliding window with heterogeneous denoisings (Local-Future Sliding Window). A dual temporal anchoring — style anchor (RoPE re-indexing) and temporal anchor (recent clean block reuse) — stabilizes the stream over unlimited durations. The student model (1.3B parameters) achieves 34 ms/frame while maintaining strong visual quality on a 400 long-duration video benchmark.

Keywords: diffusion forcing, non-verbal reactivity, synthetic DPO, 500ms latency, 6.8x speedup

Video avatarReal time
▶ Innovation: avatar reactive to interlocutor's non-verbal signals in real time — preferred in 80% of human evaluations

Avatar Forcing models user-avatar interactions through diffusion forcing, allowing the avatar to process the user's multimodal inputs (audio, head movement, expressions) in real time with approximately 500ms latency. A direct preference optimization method with synthetic losing samples enables expressive learning without additional annotations. 6.8x speedup vs baseline.

Keywords: write–manage–read, episodic memory, governance, multi-session evaluation

MemoryEvaluationGovernancePhase B
▶ Framework: write, manage, and read memory under latency, quality, and governance constraints

This 2026 survey frames agent memory as a write–manage–read loop across working, episodic, semantic, and procedural memory, evaluated over multiple sessions. It foregrounds write-time filtering, contradictions, and controlled forgetting.

Keywords: streaming diffusion, active listening, gestures, long-term consistency

Video avatarReal timePhase B
▶ Track: speaking, listening, and gestures are jointly modeled in a streaming interactive avatar

StreamAvatar adapts a video diffusion model to interactive streaming generation. It explicitly targets speaking and listening behaviors, gestures, and long-horizon consistency — beyond lip-sync or an isolated face.

Keywords: controllable voice, natural-language voice design, style cloning, multilingual

Expressive voiceEvaluationPhase B
▶ Capabilities: voice design, controllable cloning, and continuation in one model

VoxCPM2 is a technical report on an open-source voice model that brings together multilingual synthesis, description-led voice design, instruction-following style cloning, and continuation. It illustrates the shift in challenge: generating a voice is not enough; it must be possible to control how it speaks.

Very high

Keywords: STM / LTM / Knowledge layers, personalized AI agents

MemoryPhase B
▶ Key result: +37.6% accuracy vs baseline, 3-layer architecture

MemoryOS proposes a memory 'operating system' for personalized AI agents, organized in three distinct layers: short-term working memory, long-term episodic memory, and a structured knowledge base. The accuracy gains (+37.6% vs baseline) experimentally validate that memory stratification outperforms a flat context window.

Keywords: persistent memory, token reduction, conversational agents

MemoryPhase B
▶ Key results: +26% accuracy, −91% p95 latency, −90% tokens vs full-context; 90,000+ developers, 48,000+ GitHub stars

Mem0 is an open-source persistent structured memory library for AI agents. Its mechanism relies on automatic extraction of salient facts (preferences, events, context) at each conversation turn, followed by intelligent deduplication and targeted retrieval — only relevant memories are injected into the prompt at each exchange.

Keywords: temporal knowledge graph, enterprise memory, LongMemEval

MemoryPhase B
▶ Key results: +18.5% accuracy, −90% latency vs baseline on LongMemEval; outperforms MemGPT (94.8% vs 93.4%)

Zep implements Graphiti, a temporally-aware knowledge graph where each fact is timestamped and can evolve over time. The system integrates both unstructured conversational data and structured business data in a single relational graph — enabling temporal queries like 'what happened before session 3?' or 'how did the learner's understanding of this concept evolve?'.

Keywords: full-duplex, multilingual EN/FR, EU sovereignty, GDPR by design

Phase BReal time
▶ Funding: €60M (Dec. 2025) — Latency: P50 ~258ms published — Limited API access May 2026

Gradium is a production-optimized Kyutai spin-off: architecture derived from Moshi with streamed neural Mimi codec and two parallel audio streams. Multilingual EN/FR focus from the start, EU hosting, GDPR by design compliance — 'The voice infrastructure company for Europe'. Limited API access in May 2026, French WER not yet officially published.

Keywords: unified text→avatar pipeline, dual-branch diffusion transformer, zero-shot 25 FPS

Video avatarReal time
▶ Innovation: first unified end-to-end pipeline text → simultaneous speech + video in zero-shot real time

OmniTalker is the first unified end-to-end framework that simultaneously generates speech and talking head video from text and a reference video, in zero-shot real-time inference (25 FPS). Dual-branch diffusion transformer architecture: audio branch synthesizes mel-spectrograms from text, visual branch predicts poses and facial dynamics. Audio-visual fusion module ensures synchronization and stylistic coherence.

D16
Confiance Épistémique dans les Avatars Pédagogiques (IJLD, juin 2025)
Int. J. Learning and Development, vol. 15, n°2 — DOI:10.5296/ijld.v15i2.22988doi.org
Very high

Keywords: epistemic trust, learner identity, perceived competence/benevolence/integrity

Video avatar
▶ Key result: epistemic trust is the central mediator of avatar pedagogical effectiveness

Integrative theory-building review synthesizing literature on epistemic trust, learner identity formation, and organizational readiness in avatar-mediated instruction contexts in higher education. Develops a framework articulating how AI avatars influence learner engagement and reconfigure digital identity formation. Epistemic trust — belief that the avatar is a reliable source — is identified as the central mediator.

Keywords: photorealistic talking head, nuanced expressions, real-time

Video avatarReal timePhase B
▶ Key results: 40 FPS online, 512×512 resolution, talking faces with rich facial expressions from a single image

VASA-1 demonstrates that photorealistic talking faces can be generated at 40 FPS in real-time from a single image, with nuanced emotional expressions going well beyond simple lip-sync. Not yet commercialized (risk of incomplete publication), this work establishes a quality benchmark reference for expressive avatars.

Keywords: zero-shot TTS, style encoder, individual prosody, spontaneous speech

Expressive voicePhase B

PerTTS proposes a personalized and controllable zero-shot voice synthesis system: a speech style encoder captures the individual prosodic fingerprint (rhythm, intonation, spontaneous patterns) from a few audio samples, and a local prosody encoder controls phrase-by-phrase variations. The approach requires no full fine-tuning — a few minutes of audio suffice.

Keywords: disentangled facial latent space, 512×512 @ 40 FPS generation, negligible latency

Video avatarReal time
▶ Innovation: single static image + audio → 40 FPS video, independent lip/expression/gaze movements

VASA-1 generates talking avatars from a single static image and an audio clip. The disentangled facial latent space treats lip movements, expressions, gaze, and head movements as independent variables. The diffusion model operates in this latent space, enabling online generation of 512×512 videos at 40 FPS. Microsoft did not release this model publicly, acknowledging the risk of misuse.

Keywords: agentic memory, temporal knowledge graph, conversational agents

MemoryPhase B

This paper proposes an agentic memory architecture based on temporal knowledge graphs, where each graph node represents a memorized fact with its timestamp, relations to other facts, and obsolescence probability. The approach enables temporal queries on the evolution of a user's cognitive state — crucial for long-term pedagogical tracking scenarios.

Keywords: audio-driven diffusion, long-form streaming, GPU cost, open source

Video avatarReal timePhase B
▶ Reported result: 45 FPS on multi-H800 infrastructure; hardware dependencies remain decisive

LiveAvatar jointly designs algorithm and system for continuous audio-driven sequences. The repository documents substantial acceleration, but also dependence on heavy GPU infrastructure and TTS integration still listed as a next step.

Keywords: RAG, vector DB memory, conversational LLMs, survey

MemoryEvaluationPhase B

Systematic review of retrieval-augmented memory architectures (RAG) for conversational LLMs. The article synthesizes vector database approaches, compares chunking and indexing strategies, and evaluates trade-offs between retrieval latency and contextual accuracy.

Keywords: RAG, long-term memory, agentic agents, reinforcement learning

MemoryPhase B

This paper documents the transition from static RAG architectures to agentic long-term memory systems, where the agent learns to manage its own memory via reinforcement learning. The approach enables active selection of memories to retain or forget, rather than simple mechanical compression.

Keywords: complete digital human, 3D avatar, expressive speech, grounded dialogue

Expressive voicePhase B

Hi-Reco proposes a rare integrated approach: a complete digital human combining 3D avatar, expressive speech, and grounded dialogue in a unified architecture. Unlike modular pipelines (separate TTS + avatar), the end-to-end approach guarantees coherence between vocal and facial expression.

Keywords: talking head synthesis, comprehensive survey, real-time / expressiveness / quality trilemma

Video avatarReal timeEvaluationPhase B

Comprehensive review of talking head synthesis techniques published in ACM Computing Surveys (2025). The article formally documents the fundamental trilemma: no current solution simultaneously satisfies all three constraints — real-time (<500ms), rich emotional expressiveness, and photorealistic quality. This trilemma structures the research space.

Keywords: TTS, complex style control, benchmark, evaluation 11Labs / Deepgram / OpenAI

Expressive voiceEvaluationPhase B

EmergentTTS-Eval is a NeurIPS 2025 benchmark for evaluating complex style control in voice synthesis. It evaluates 11 expressive dimensions (emotion, intensity, rhythm, character) on major market TTS systems (ElevenLabs, Deepgram, OpenAI 4o-mini-TTS) and establishes reproducible comparative metrics.

Keywords: multimodal speech model, contextual prosody, natural backchannels, RVQ

Expressive voicePhase B
▶ License: Apache 2.0 — 1B parameters — self-hosted on GPU (A100 recommended)

Sesame CSM is a multimodal conversational speech model designed to 'cross the uncanny valley of conversational voice': it generates RVQ audio codes from text and audio inputs, with contextual prosody, natural backchannels ('mm-hm', 'yes') and human turn-taking behavior. Not optimized for production streaming but a high-quality research reference.

Keywords: tandem architecture, agentic memory, knowledge graph, real-time RAG

MemoryReal timePhase B

The KAME Tandem architecture combines a dynamic knowledge graph with real-time agentic memory. Its 'tandem' approach couples two complementary components: a fast retrieval module for immediate facts, and a deep reasoning module for complex semantic relationships. This is the Slow component of the MoshiRAG architecture in GamiWays taxonomy.

Keywords: Timestep-forcing Pipeline Parallelism, Rolling Sink Frame, identity drift, 14B parameters

Video avatarReal time
▶ Innovation: 20 FPS end-to-end on 5 H800 GPUs, elimination of identity drift on long sequences

Live Avatar resolves the video diffusion generation bottleneck for real-time applications. Timestep-forcing Pipeline Parallelism (TPP) pipelines denoising steps across multiple GPUs. The Rolling Sink Frame Mechanism (RSFM) recalibrates appearance from a cached reference image, eliminating identity drift on long sequences. 20 FPS end-to-end on 5 H800 GPUs with 14B parameters.

Keywords: Anchor-Heavy Identity Sinks, 20x inference cost reduction, multi-turn, outperforms Sora2/Veo3

Video avatarReal time
▶ Key result: outperforms Sora2 and Veo3 in multi-turn video coherence, latency from 1-2 min → real time

LiveTalk combines a video diffusion model conditioned on text, image, and audio with audio language models. Anchor-Heavy Identity Sinks technique for long-duration inference. The distilled model achieves comparable quality to bidirectional baselines with 20x less inference cost. In multi-turn evaluation, LiveTalk outperforms Sora2 and Veo3 in video coherence, bringing latency from 1-2 minutes to real time.

Keywords: comprehensive survey, GANs→diffusion transformers, HDTF/VFHQ/LRS3 benchmarks, ethnic diversity

Video avatarEvaluation

Systematic review covering the evolution of talking head generation techniques from GANs (Wav2Lip, SadTalker) to recent diffusion transformers. Maps standard benchmarks (HDTF, VFHQ, LRS3), evaluation metrics (FID, FVD, SyncNet, PSNR), and architectural trade-offs. Identifies four priority improvement axes: identity on long sequences, expressiveness beyond lips, ethnic and cultural diversity, inference speed on constrained hardware.

D12
Fidélité Visuelle, Présence et Réponses Physiologiques (CHB, 2025)
Computers in Human Behavior — DOI:10.1016/j.chb.2025.108596doi.org
High

Keywords: HRV/EEG, sympathetic stress, presence-anonymity trilemma, low/medium/high visual fidelity

Video avatar
▶ Key result: high-fidelity avatars → higher sympathetic stress during sensitive conversations

Experimental study measuring physiological (HRV, EEG) and psychological (perceived stress) responses of participants interacting with low, medium, and high visual fidelity avatars. High-fidelity avatars generate a significantly higher LF/HF ratio (sympathetic stress) than low-fidelity avatars, particularly during sensitive conversations. Confirms the presence-anonymity trilemma.

Keywords: parasocial interaction, digital attachment, emotional dependency, ethical design

Video avatarGovernance

Systematic review of two decades of research proposing a theoretical model of the evolution of human-AI emotional relationships: from parasocial interaction (unilateral emotional investment) to digital attachment. Identified determinants: AI personalization, memory continuity, perceived emotional reactivity, anthropomorphization. Dependency risks and ethical design criteria proposed.

D14
Pseudo-intimité Émotionnelle et IA (Frontiers in Psychology, sept. 2025)
Frontiers in Psychology — DOI:10.3389/fpsyg.2025.1679324doi.org
High

Keywords: pseudo-intimacy, affective AI, anthropomorphization bias, transparency, informed consent

Video avatar

Explores the risk of pseudo-intimacy with affective AIs: emotional AIs create relationships that mimic intimacy without the conditions of reciprocity, shared vulnerability, and common history that found authentic human relationships. Analyzes cognitive mechanisms (anthropomorphization bias, affective reinforcement by simulated reactivity) and proposes an affective design ethics framework based on transparency, deliberate expressiveness limitation, and informed consent.

D15
Avatars dans le Métaverse Éducatif — Revue Systématique (PMC, juin 2025)
Visual Computing for Industry, Biomedicine, and Art — DOI:10.1186/s42492-025-00196-9doi.org
High

Keywords: educational avatars, metaverse, generative NLP, data security, algorithmic bias

Video avatarGovernance

Systematic review evaluating avatar contributions to learning in immersive virtual environments. Documents that avatars enable individualized experiences and improved collaborative activities, and that advances in NLP and generative models have significantly improved their capabilities. Identifies major obstacles: data security, ethical concerns (algorithmic bias, privacy), and limited infrastructure.

Keywords: ethnic bias, DH-FaceVid-1K dataset, multilingual generalization, algorithmic fairness, cultural representation

Video avatarGovernance
▶ Documented gap: current avatar generation models are biased toward European and American ethnicities — degraded performance for non-European users

This ACM Grand Challenge directly targets the major gap in talking head generation models: their poor generalization on non-European and non-American ethnicities. The DH-FaceVid-1K dataset is introduced to address this gap. Participants must develop models integrating diffusion and GAN techniques to improve lip synchronization, identity coherence, and inference speed on multilingual and multi-ethnic data. This challenge reveals that standard benchmarks (HDTF, VFHQ) are massively biased toward Caucasian English-speaking faces.

Keywords: long-term memory benchmark, LLM assistants, systematic evaluation

MemoryEvaluationPhase B

This benchmark establishes a systematic evaluation framework for long-term memory capabilities of LLM assistants, covering fact retention, temporal coherence, and progressive personalization. It documents current limits of naive approaches (full context, truncation) and opens the path toward genuinely personalized long-term assistants.

Keywords: context compression, LLM inference, long sessions, 128K tokens

MemoryPhase B

This research proposes stateful context compression enabling a 128K token attention window without perceptible quality degradation. Unlike brutal truncation, the stateful approach preserves a compressed representation of history — particularly applicable to long conversational sessions (1h+).

Keywords: digital twin ownership, data autonomy, renewed social contract, post-mortem rights

Video avatarGovernance

Argues that natural persons must be recognized as moral and legal owners of their AI digital twins, as these entities are intimate extensions of the individual built from their personal data. Current legal frameworks, which privilege technological infrastructure over data autonomy, are criticized as perpetuating systemic inadequacies. Proposes a renewed social contract centered on individual dominion and consent.

D11
La Vallée de l'Étrange Audio-Visuelle (ACM CHI 2022)
ACM CHI 2022 — DOI:10.1145/3491102.3517564dl.acm.org
High

Keywords: vocal uncanny valley, neural TTS vs human voice, trust, speaker/listener gender

Video avatar
▶ Key result: neural TTS → significantly less trustworthy than real human voice — effect modulated by gender

This ACM CHI study demonstrates that a new dimension of the uncanny valley emerges in the vocal domain: a virtual human using neural TTS is perceived as significantly less trustworthy than a virtual human using a real human voice. The effect is modulated by speaker and listener gender. The study extends Mori's (1970) classic concept from visual morphology to acoustic fidelity.

Keywords: Audio-Pose Prior Refocusing, rhythm-gesture coupling, bidirectional cross-attention

Video avatar

ApoAvatar introduces an Audio-Pose Prior Refocusing mechanism that links speech style to movement dynamics: strong accents amplify gesture amplitude, silent passages suppress superfluous movements. A frame-wise audio-video interaction module updates audio features with current visual context via bidirectional cross-attention. Clear gains on lip synchronization, gesture expressiveness, and overall naturalness.

Keywords: speech-synchronized hand gestures, audio-driven poses, perceived authenticity

Video avatar

EMO2 extends audio-driven avatar generation beyond the face to include speech-synchronized hand gestures. Two-step process: hand pose generation directly from audio (strong audio-gesture correlation), then diffusion model synthesizing video frames integrating these poses. Outperforms CyberHost and Vlogger on visual quality and synchronization accuracy.

Keywords: EU AI Act Art. 50(2), mandatory labeling, €35M penalties, human rights, prohibited emotional recognition

Video avatarGovernance
▶ Legal obligation from August 2026: all generated video avatars must be clearly and visibly labeled as AI content

Analyzes how the EU AI Act (adopted 2024, full enforcement August 2026) establishes a rights-centered framework for regulating synthetic media. Article 50(2) requires that generative outputs including deepfakes be clearly and visibly labeled. Penalties: up to €35M or 7% of global revenue. Emotional recognition in educational settings is explicitly prohibited by Article 5 since February 2025.

Keywords: AI inference carbon footprint, 0.24 Wh/Gemini query, video generation 10–100× more costly, IDC 23 TWh

Video avatarEvaluationGovernance
▶ Key data: median Gemini query = 0.24 Wh / 0.03 g CO₂; large open-source models (70B) = up to 1.7 Wh/query

MIT Technology Review's analysis evaluates AI energy footprint at sector scale: IDC estimates AI data centers consumed 23 TWh in 2022. Google publishes a 2025 methodology for measuring AI inference environmental impact: a median Gemini query consumes 0.24 Wh and emits 0.03 g CO₂. Large open-source models (70B parameters) consume up to 1.7 Wh per query. Video generation is significantly more costly than text generation — likely 10 to 100 times more energy-intensive for real-time avatar streaming.

Browse the detailed domain-by-domain classification

This view keeps the corpus’ historical organization. Use the filters above first to isolate a decision, then return here to read references by their original academic domain.

Domain A — Conversational Memory

7 publications covering memory stratification, context compression, RAG architectures, and temporal knowledge graphs. Directly applicable to Core Engine hypothesis H2.

Very high

Keywords: STM / LTM / Knowledge layers, personalized AI agents

▶ Key result: +37.6% accuracy vs baseline, 3-layer architecture

MemoryOS proposes a memory 'operating system' for personalized AI agents, organized in three distinct layers: short-term working memory, long-term episodic memory, and a structured knowledge base. The accuracy gains (+37.6% vs baseline) experimentally validate that memory stratification outperforms a flat context window.

Keywords: persistent memory, token reduction, conversational agents

▶ Key results: +26% accuracy, −91% p95 latency, −90% tokens vs full-context; 90,000+ developers, 48,000+ GitHub stars

Mem0 is an open-source persistent structured memory library for AI agents. Its mechanism relies on automatic extraction of salient facts (preferences, events, context) at each conversation turn, followed by intelligent deduplication and targeted retrieval — only relevant memories are injected into the prompt at each exchange.

Keywords: long-term memory benchmark, LLM assistants, systematic evaluation

This benchmark establishes a systematic evaluation framework for long-term memory capabilities of LLM assistants, covering fact retention, temporal coherence, and progressive personalization. It documents current limits of naive approaches (full context, truncation) and opens the path toward genuinely personalized long-term assistants.

Keywords: context compression, LLM inference, long sessions, 128K tokens

This research proposes stateful context compression enabling a 128K token attention window without perceptible quality degradation. Unlike brutal truncation, the stateful approach preserves a compressed representation of history — particularly applicable to long conversational sessions (1h+).

Keywords: RAG, vector DB memory, conversational LLMs, survey

Systematic review of retrieval-augmented memory architectures (RAG) for conversational LLMs. The article synthesizes vector database approaches, compares chunking and indexing strategies, and evaluates trade-offs between retrieval latency and contextual accuracy.

Keywords: RAG, long-term memory, agentic agents, reinforcement learning

This paper documents the transition from static RAG architectures to agentic long-term memory systems, where the agent learns to manage its own memory via reinforcement learning. The approach enables active selection of memories to retain or forget, rather than simple mechanical compression.

Keywords: temporal knowledge graph, enterprise memory, LongMemEval

▶ Key results: +18.5% accuracy, −90% latency vs baseline on LongMemEval; outperforms MemGPT (94.8% vs 93.4%)

Zep implements Graphiti, a temporally-aware knowledge graph where each fact is timestamped and can evolve over time. The system integrates both unstructured conversational data and structured business data in a single relational graph — enabling temporal queries like 'what happened before session 3?' or 'how did the learner's understanding of this concept evolve?'.

Domain B — Avatar & Voice Synthesis

8 publications covering photorealistic avatars, personalized voice synthesis, TTS benchmarks, and multimodal speech models. Directly applicable to R&D axes 1, 2a, and 3.

Keywords: photorealistic talking head, nuanced expressions, real-time

▶ Key results: 40 FPS online, 512×512 resolution, talking faces with rich facial expressions from a single image

VASA-1 demonstrates that photorealistic talking faces can be generated at 40 FPS in real-time from a single image, with nuanced emotional expressions going well beyond simple lip-sync. Not yet commercialized (risk of incomplete publication), this work establishes a quality benchmark reference for expressive avatars.

Keywords: audio-avatar LLM, expressive facial movements, LoRA architecture

▶ Architecture: 8B parameters + 0.16B LoRA

A²-LLM is an end-to-end audio-avatar LLM that generates emotionally rich facial movements beyond lip-sync, by directly coupling text generation and facial synthesis in a single model. The 8B + 0.16B LoRA architecture enables fine adaptation to different expressive styles without retraining the entire model.

Keywords: complete digital human, 3D avatar, expressive speech, grounded dialogue

Hi-Reco proposes a rare integrated approach: a complete digital human combining 3D avatar, expressive speech, and grounded dialogue in a unified architecture. Unlike modular pipelines (separate TTS + avatar), the end-to-end approach guarantees coherence between vocal and facial expression.

Keywords: talking head synthesis, comprehensive survey, real-time / expressiveness / quality trilemma

Comprehensive review of talking head synthesis techniques published in ACM Computing Surveys (2025). The article formally documents the fundamental trilemma: no current solution simultaneously satisfies all three constraints — real-time (<500ms), rich emotional expressiveness, and photorealistic quality. This trilemma structures the research space.

Keywords: streaming diffusion, Local-Future Sliding-Window Denoising, long-form real-time

▶ Innovation: single image + streaming audio → real-time long-form video in 1 diffusion step

AvatarForcing is the most directly applicable academic paper to R&D Axis 1 (latency reduction). Its 'Local-Future Sliding-Window Denoising' technique enables talking avatar video generation in 1-step streaming diffusion — drastically reducing denoising iterations (from 20–50 steps to 1 step), dividing generation time by 60–90%.

Keywords: TTS, complex style control, benchmark, evaluation 11Labs / Deepgram / OpenAI

EmergentTTS-Eval is a NeurIPS 2025 benchmark for evaluating complex style control in voice synthesis. It evaluates 11 expressive dimensions (emotion, intensity, rhythm, character) on major market TTS systems (ElevenLabs, Deepgram, OpenAI 4o-mini-TTS) and establishes reproducible comparative metrics.

Keywords: zero-shot TTS, style encoder, individual prosody, spontaneous speech

PerTTS proposes a personalized and controllable zero-shot voice synthesis system: a speech style encoder captures the individual prosodic fingerprint (rhythm, intonation, spontaneous patterns) from a few audio samples, and a local prosody encoder controls phrase-by-phrase variations. The approach requires no full fine-tuning — a few minutes of audio suffice.

Keywords: multimodal speech model, contextual prosody, natural backchannels, RVQ

▶ License: Apache 2.0 — 1B parameters — self-hosted on GPU (A100 recommended)

Sesame CSM is a multimodal conversational speech model designed to 'cross the uncanny valley of conversational voice': it generates RVQ audio codes from text and audio inputs, with contextual prosody, natural backchannels ('mm-hm', 'yes') and human turn-taking behavior. Not optimized for production streaming but a high-quality research reference.

Future Perspectives 2026–2028 — Full-Duplex Architecture & Agentic Memory

6 recent works (2025–2026) documenting full-duplex interaction models, agentic memory architectures, and European sovereignty solutions. These works inform GamiWays' Phase B/C architectural choices.

Architecture: MoE 276B/12B active, native dMel audio tokenization, dual Fast/Slow model

▶ Status: Waitlist preview, May 2026 — no public API

TML Interaction-Small is a MoE model with 276B total parameters but only 12B active per inference. It integrates native dMel audio tokenization and a dual architecture: the Fast Model (200ms chunks) handles real-time conversation while the Slow Model (asynchronous) handles tool calls, RAG, and deep reasoning. Visual proactivity and absolute temporal awareness. Known limitation: coherence degrading after ~30–45 minutes.

Keywords: full-duplex, asynchronous RAG, Moshi, enriched memory

▶ Target latency: 160–260ms + contextual enrichment

MoshiRAG is an extension of Moshi (Kyutai) integrating asynchronous RAG and enriched memory in a full-duplex model. The target architecture is exactly GamiWays' Phase C: full-duplex + RAG/retrieval/reasoning in parallel, without blocking the audio stream. The GamiWays Core Engine is architecturally this orchestration layer.

Keywords: tandem architecture, agentic memory, knowledge graph, real-time RAG

The KAME Tandem architecture combines a dynamic knowledge graph with real-time agentic memory. Its 'tandem' approach couples two complementary components: a fast retrieval module for immediate facts, and a deep reasoning module for complex semantic relationships. This is the Slow component of the MoshiRAG architecture in GamiWays taxonomy.

Keywords: duplex cascade, conversational pipeline, reduced latency, streaming

DuplexCascade proposes a duplex cascade conversational pipeline architecture, reducing perceived latency by starting each step (TTS, avatar generation) from the first tokens of the previous step, without waiting for completion. This is the academic formalization of the 'streaming pipeline' approach GamiWays targets for Phase A+ (estimated gain: −40% total latency).

Keywords: agentic memory, temporal knowledge graph, conversational agents

This paper proposes an agentic memory architecture based on temporal knowledge graphs, where each graph node represents a memorized fact with its timestamp, relations to other facts, and obsolescence probability. The approach enables temporal queries on the evolution of a user's cognitive state — crucial for long-term pedagogical tracking scenarios.

Keywords: full-duplex, multilingual EN/FR, EU sovereignty, GDPR by design

▶ Funding: €60M (Dec. 2025) — Latency: P50 ~258ms published — Limited API access May 2026

Gradium is a production-optimized Kyutai spin-off: architecture derived from Moshi with streamed neural Mimi codec and two parallel audio streams. Multilingual EN/FR focus from the start, EU hosting, GDPR by design compliance — 'The voice infrastructure company for Europe'. Limited API access in May 2026, French WER not yet officially published.

Domain D — Video Streaming Avatars: Technical, Psychology, Law & Ethics

20 references covering 6 dimensions: technical foundations (VASA-1, RETA, AvatarForcing), psychology (uncanny valley, attachment, epistemic trust), social and pedagogical implications, law and regulation (EU AI Act, biometric data, avatar ownership), environmental impact and sustainability, and cross-cutting critical perspectives (ethnic bias).

Part I — Technical Foundations

Keywords: disentangled facial latent space, 512×512 @ 40 FPS generation, negligible latency

▶ Innovation: single static image + audio → 40 FPS video, independent lip/expression/gaze movements

VASA-1 generates talking avatars from a single static image and an audio clip. The disentangled facial latent space treats lip movements, expressions, gaze, and head movements as independent variables. The diffusion model operates in this latent space, enabling online generation of 512×512 videos at 40 FPS. Microsoft did not release this model publicly, acknowledging the risk of misuse.

Keywords: real-time / sync / emotion trilemma, 3DMM, single-pass GAN, 55+ FPS

▶ Innovation: first framework to simultaneously resolve all 3 constraints — 55+ FPS, lip sync, label-free emotional fidelity

RETA is the first framework to simultaneously resolve the fundamental trilemma: real-time performance, lip synchronization accuracy, and emotional fidelity. It disentangles the audio signal into two representations: a 3DMM geometry for lip synchronization, and a dynamic emotional embedding learned without labels via cross-modal distillation. A single-pass GAN generator integrates these representations hierarchically. 55+ FPS with new SOTA on all metrics.

Keywords: Timestep-forcing Pipeline Parallelism, Rolling Sink Frame, identity drift, 14B parameters

▶ Innovation: 20 FPS end-to-end on 5 H800 GPUs, elimination of identity drift on long sequences

Live Avatar resolves the video diffusion generation bottleneck for real-time applications. Timestep-forcing Pipeline Parallelism (TPP) pipelines denoising steps across multiple GPUs. The Rolling Sink Frame Mechanism (RSFM) recalibrates appearance from a cached reference image, eliminating identity drift on long sequences. 20 FPS end-to-end on 5 H800 GPUs with 14B parameters.

Keywords: Local-Future Sliding-Window, 1.3B parameters, 34 ms/frame, 2-step distillation

▶ Innovation: 34 ms/frame (≈29 FPS), 1.3B parameters — lightest and most actionable architecture for production

AvatarForcing resolves the exposure bias of autoregressive models via a sliding window with heterogeneous denoisings (Local-Future Sliding Window). A dual temporal anchoring — style anchor (RoPE re-indexing) and temporal anchor (recent clean block reuse) — stabilizes the stream over unlimited durations. The student model (1.3B parameters) achieves 34 ms/frame while maintaining strong visual quality on a 400 long-duration video benchmark.

Keywords: diffusion forcing, non-verbal reactivity, synthetic DPO, 500ms latency, 6.8x speedup

▶ Innovation: avatar reactive to interlocutor's non-verbal signals in real time — preferred in 80% of human evaluations

Avatar Forcing models user-avatar interactions through diffusion forcing, allowing the avatar to process the user's multimodal inputs (audio, head movement, expressions) in real time with approximately 500ms latency. A direct preference optimization method with synthetic losing samples enables expressive learning without additional annotations. 6.8x speedup vs baseline.

Keywords: unified text→avatar pipeline, dual-branch diffusion transformer, zero-shot 25 FPS

▶ Innovation: first unified end-to-end pipeline text → simultaneous speech + video in zero-shot real time

OmniTalker is the first unified end-to-end framework that simultaneously generates speech and talking head video from text and a reference video, in zero-shot real-time inference (25 FPS). Dual-branch diffusion transformer architecture: audio branch synthesizes mel-spectrograms from text, visual branch predicts poses and facial dynamics. Audio-visual fusion module ensures synchronization and stylistic coherence.

Keywords: speech-synchronized hand gestures, audio-driven poses, perceived authenticity

EMO2 extends audio-driven avatar generation beyond the face to include speech-synchronized hand gestures. Two-step process: hand pose generation directly from audio (strong audio-gesture correlation), then diffusion model synthesizing video frames integrating these poses. Outperforms CyberHost and Vlogger on visual quality and synchronization accuracy.

Keywords: Audio-Pose Prior Refocusing, rhythm-gesture coupling, bidirectional cross-attention

ApoAvatar introduces an Audio-Pose Prior Refocusing mechanism that links speech style to movement dynamics: strong accents amplify gesture amplitude, silent passages suppress superfluous movements. A frame-wise audio-video interaction module updates audio features with current visual context via bidirectional cross-attention. Clear gains on lip synchronization, gesture expressiveness, and overall naturalness.

Keywords: Anchor-Heavy Identity Sinks, 20x inference cost reduction, multi-turn, outperforms Sora2/Veo3

▶ Key result: outperforms Sora2 and Veo3 in multi-turn video coherence, latency from 1-2 min → real time

LiveTalk combines a video diffusion model conditioned on text, image, and audio with audio language models. Anchor-Heavy Identity Sinks technique for long-duration inference. The distilled model achieves comparable quality to bidirectional baselines with 20x less inference cost. In multi-turn evaluation, LiveTalk outperforms Sora2 and Veo3 in video coherence, bringing latency from 1-2 minutes to real time.

Keywords: comprehensive survey, GANs→diffusion transformers, HDTF/VFHQ/LRS3 benchmarks, ethnic diversity

Systematic review covering the evolution of talking head generation techniques from GANs (Wav2Lip, SadTalker) to recent diffusion transformers. Maps standard benchmarks (HDTF, VFHQ, LRS3), evaluation metrics (FID, FVD, SyncNet, PSNR), and architectural trade-offs. Identifies four priority improvement axes: identity on long sequences, expressiveness beyond lips, ethnic and cultural diversity, inference speed on constrained hardware.

Part II — Psychology and Behavior
D11
La Vallée de l'Étrange Audio-Visuelle (ACM CHI 2022)
ACM CHI 2022 — DOI:10.1145/3491102.3517564dl.acm.org
High

Keywords: vocal uncanny valley, neural TTS vs human voice, trust, speaker/listener gender

▶ Key result: neural TTS → significantly less trustworthy than real human voice — effect modulated by gender

This ACM CHI study demonstrates that a new dimension of the uncanny valley emerges in the vocal domain: a virtual human using neural TTS is perceived as significantly less trustworthy than a virtual human using a real human voice. The effect is modulated by speaker and listener gender. The study extends Mori's (1970) classic concept from visual morphology to acoustic fidelity.

D12
Fidélité Visuelle, Présence et Réponses Physiologiques (CHB, 2025)
Computers in Human Behavior — DOI:10.1016/j.chb.2025.108596doi.org
High

Keywords: HRV/EEG, sympathetic stress, presence-anonymity trilemma, low/medium/high visual fidelity

▶ Key result: high-fidelity avatars → higher sympathetic stress during sensitive conversations

Experimental study measuring physiological (HRV, EEG) and psychological (perceived stress) responses of participants interacting with low, medium, and high visual fidelity avatars. High-fidelity avatars generate a significantly higher LF/HF ratio (sympathetic stress) than low-fidelity avatars, particularly during sensitive conversations. Confirms the presence-anonymity trilemma.

Keywords: parasocial interaction, digital attachment, emotional dependency, ethical design

Systematic review of two decades of research proposing a theoretical model of the evolution of human-AI emotional relationships: from parasocial interaction (unilateral emotional investment) to digital attachment. Identified determinants: AI personalization, memory continuity, perceived emotional reactivity, anthropomorphization. Dependency risks and ethical design criteria proposed.

D14
Pseudo-intimité Émotionnelle et IA (Frontiers in Psychology, sept. 2025)
Frontiers in Psychology — DOI:10.3389/fpsyg.2025.1679324doi.org
High

Keywords: pseudo-intimacy, affective AI, anthropomorphization bias, transparency, informed consent

Explores the risk of pseudo-intimacy with affective AIs: emotional AIs create relationships that mimic intimacy without the conditions of reciprocity, shared vulnerability, and common history that found authentic human relationships. Analyzes cognitive mechanisms (anthropomorphization bias, affective reinforcement by simulated reactivity) and proposes an affective design ethics framework based on transparency, deliberate expressiveness limitation, and informed consent.

Part III — Social and Pedagogical Implications
D15
Avatars dans le Métaverse Éducatif — Revue Systématique (PMC, juin 2025)
Visual Computing for Industry, Biomedicine, and Art — DOI:10.1186/s42492-025-00196-9doi.org
High

Keywords: educational avatars, metaverse, generative NLP, data security, algorithmic bias

Systematic review evaluating avatar contributions to learning in immersive virtual environments. Documents that avatars enable individualized experiences and improved collaborative activities, and that advances in NLP and generative models have significantly improved their capabilities. Identifies major obstacles: data security, ethical concerns (algorithmic bias, privacy), and limited infrastructure.

D16
Confiance Épistémique dans les Avatars Pédagogiques (IJLD, juin 2025)
Int. J. Learning and Development, vol. 15, n°2 — DOI:10.5296/ijld.v15i2.22988doi.org
Very high

Keywords: epistemic trust, learner identity, perceived competence/benevolence/integrity

▶ Key result: epistemic trust is the central mediator of avatar pedagogical effectiveness

Integrative theory-building review synthesizing literature on epistemic trust, learner identity formation, and organizational readiness in avatar-mediated instruction contexts in higher education. Develops a framework articulating how AI avatars influence learner engagement and reconfigure digital identity formation. Epistemic trust — belief that the avatar is a reliable source — is identified as the central mediator.

Part IV — Law, Data and Regulation

Keywords: EU AI Act Art. 50(2), mandatory labeling, €35M penalties, human rights, prohibited emotional recognition

▶ Legal obligation from August 2026: all generated video avatars must be clearly and visibly labeled as AI content

Analyzes how the EU AI Act (adopted 2024, full enforcement August 2026) establishes a rights-centered framework for regulating synthetic media. Article 50(2) requires that generative outputs including deepfakes be clearly and visibly labeled. Penalties: up to €35M or 7% of global revenue. Emotional recognition in educational settings is explicitly prohibited by Article 5 since February 2025.

Keywords: digital twin ownership, data autonomy, renewed social contract, post-mortem rights

Argues that natural persons must be recognized as moral and legal owners of their AI digital twins, as these entities are intimate extensions of the individual built from their personal data. Current legal frameworks, which privilege technological infrastructure over data autonomy, are criticized as perpetuating systemic inadequacies. Proposes a renewed social contract centered on individual dominion and consent.

Part V — Environmental Impact and Sustainability

Keywords: AI inference carbon footprint, 0.24 Wh/Gemini query, video generation 10–100× more costly, IDC 23 TWh

▶ Key data: median Gemini query = 0.24 Wh / 0.03 g CO₂; large open-source models (70B) = up to 1.7 Wh/query

MIT Technology Review's analysis evaluates AI energy footprint at sector scale: IDC estimates AI data centers consumed 23 TWh in 2022. Google publishes a 2025 methodology for measuring AI inference environmental impact: a median Gemini query consumes 0.24 Wh and emits 0.03 g CO₂. Large open-source models (70B parameters) consume up to 1.7 Wh per query. Video generation is significantly more costly than text generation — likely 10 to 100 times more energy-intensive for real-time avatar streaming.

Part VI — Cross-Cutting Critical Perspectives

Keywords: ethnic bias, DH-FaceVid-1K dataset, multilingual generalization, algorithmic fairness, cultural representation

▶ Documented gap: current avatar generation models are biased toward European and American ethnicities — degraded performance for non-European users

This ACM Grand Challenge directly targets the major gap in talking head generation models: their poor generalization on non-European and non-American ethnicities. The DH-FaceVid-1K dataset is introduced to address this gap. Participants must develop models integrating diffusion and GAN techniques to improve lip synchronization, identity coherence, and inference speed on multilingual and multi-ethnic data. This challenge reveals that standard benchmarks (HDTF, VFHQ) are massively biased toward Caucasian English-speaking faces.

Summary — Maturity levels by domain

Source: gamiways-recherches-2026.md — May 2026

DomainAcademic maturityCommercial availabilityGamiWays gap
Conversational memoryHigh (Mem0, MemoryOS)Partial (Mem0 API)3-layer architecture + avatar-specific SLM distillation
Multi-session avatar integrationHigh (VASA-1, A²-LLM)Partial (HeyGen, Tavus)Latency <500ms + full body
Personalized expressive TTSHigh (PerTTS, EmergentTTS)Good (ElevenLabs, Cartesia)Individual prosodic fingerprint
Conversational orchestrationEmerging (MoshiRAG, KAME)LowConfigurable freedom spectrum + agentic Game Master

Open Research Questions

These unresolved questions structure the GamiWays R&D program 2026–2028.