Research decisions
SectionDirect relevance
What research unlocks right now
Four Phase B decisions to inform before multiplying integrations or demonstrators.
Measure useful memory
Separate what must be retained, summarized, contradicted, or forgotten across sessions — rather than merely extending context.
Control how, not only what
Evaluate intent, pace, silence, and prosody scene by scene so characters do not merely recite a response.
Test video presence
Measure speaking, listening, non-verbal response, and identity stability separately instead of scoring the avatar globally.
Arbitrate cost, sovereignty, and trust
Compare architectures on perceived latency, infrastructure requirements, data rights, and the ability to fall back to audio.
Strategic synthesis
SectionDecision matrix
Prioritize what must be validated, decided, or secured
Choose a dimension, then narrow the list by impact or context. Each issue opens a research reading, a priority action, and the next paths to pursue.
Filter by impact
Filter by context
6 issues shown
Navigation — Related pages
Corpus & protocols
SectionFind research by the decision it can inform
Filter references by the issue at hand instead of scanning a linear bibliography. Each card connects a result to a concrete hypothesis, test, or trade-off for the Core, Storygami, or Edugami.
45 references shown
Keywords: audio-avatar LLM, expressive facial movements, LoRA architecture
A²-LLM is an end-to-end audio-avatar LLM that generates emotionally rich facial movements beyond lip-sync, by directly coupling text generation and facial synthesis in a single model. The 8B + 0.16B LoRA architecture enables fine adaptation to different expressive styles without retraining the entire model.
Keywords: streaming diffusion, Local-Future Sliding-Window Denoising, long-form real-time
AvatarForcing is the most directly applicable academic paper to R&D Axis 1 (latency reduction). Its 'Local-Future Sliding-Window Denoising' technique enables talking avatar video generation in 1-step streaming diffusion — drastically reducing denoising iterations (from 20–50 steps to 1 step), dividing generation time by 60–90%.
Architecture: MoE 276B/12B active, native dMel audio tokenization, dual Fast/Slow model
TML Interaction-Small is a MoE model with 276B total parameters but only 12B active per inference. It integrates native dMel audio tokenization and a dual architecture: the Fast Model (200ms chunks) handles real-time conversation while the Slow Model (asynchronous) handles tool calls, RAG, and deep reasoning. Visual proactivity and absolute temporal awareness. Known limitation: coherence degrading after ~30–45 minutes.
Keywords: full-duplex, asynchronous RAG, Moshi, enriched memory
MoshiRAG is an extension of Moshi (Kyutai) integrating asynchronous RAG and enriched memory in a full-duplex model. The target architecture is exactly GamiWays' Phase C: full-duplex + RAG/retrieval/reasoning in parallel, without blocking the audio stream. The GamiWays Core Engine is architecturally this orchestration layer.
Keywords: duplex cascade, conversational pipeline, reduced latency, streaming
DuplexCascade proposes a duplex cascade conversational pipeline architecture, reducing perceived latency by starting each step (TTS, avatar generation) from the first tokens of the previous step, without waiting for completion. This is the academic formalization of the 'streaming pipeline' approach GamiWays targets for Phase A+ (estimated gain: −40% total latency).
Keywords: real-time / sync / emotion trilemma, 3DMM, single-pass GAN, 55+ FPS
RETA is the first framework to simultaneously resolve the fundamental trilemma: real-time performance, lip synchronization accuracy, and emotional fidelity. It disentangles the audio signal into two representations: a 3DMM geometry for lip synchronization, and a dynamic emotional embedding learned without labels via cross-modal distillation. A single-pass GAN generator integrates these representations hierarchically. 55+ FPS with new SOTA on all metrics.
Keywords: Local-Future Sliding-Window, 1.3B parameters, 34 ms/frame, 2-step distillation
AvatarForcing resolves the exposure bias of autoregressive models via a sliding window with heterogeneous denoisings (Local-Future Sliding Window). A dual temporal anchoring — style anchor (RoPE re-indexing) and temporal anchor (recent clean block reuse) — stabilizes the stream over unlimited durations. The student model (1.3B parameters) achieves 34 ms/frame while maintaining strong visual quality on a 400 long-duration video benchmark.
Keywords: diffusion forcing, non-verbal reactivity, synthetic DPO, 500ms latency, 6.8x speedup
Avatar Forcing models user-avatar interactions through diffusion forcing, allowing the avatar to process the user's multimodal inputs (audio, head movement, expressions) in real time with approximately 500ms latency. A direct preference optimization method with synthetic losing samples enables expressive learning without additional annotations. 6.8x speedup vs baseline.
Keywords: write–manage–read, episodic memory, governance, multi-session evaluation
This 2026 survey frames agent memory as a write–manage–read loop across working, episodic, semantic, and procedural memory, evaluated over multiple sessions. It foregrounds write-time filtering, contradictions, and controlled forgetting.
Keywords: streaming diffusion, active listening, gestures, long-term consistency
StreamAvatar adapts a video diffusion model to interactive streaming generation. It explicitly targets speaking and listening behaviors, gestures, and long-horizon consistency — beyond lip-sync or an isolated face.
Keywords: controllable voice, natural-language voice design, style cloning, multilingual
VoxCPM2 is a technical report on an open-source voice model that brings together multilingual synthesis, description-led voice design, instruction-following style cloning, and continuation. It illustrates the shift in challenge: generating a voice is not enough; it must be possible to control how it speaks.
Keywords: STM / LTM / Knowledge layers, personalized AI agents
MemoryOS proposes a memory 'operating system' for personalized AI agents, organized in three distinct layers: short-term working memory, long-term episodic memory, and a structured knowledge base. The accuracy gains (+37.6% vs baseline) experimentally validate that memory stratification outperforms a flat context window.
Keywords: persistent memory, token reduction, conversational agents
Mem0 is an open-source persistent structured memory library for AI agents. Its mechanism relies on automatic extraction of salient facts (preferences, events, context) at each conversation turn, followed by intelligent deduplication and targeted retrieval — only relevant memories are injected into the prompt at each exchange.
Keywords: temporal knowledge graph, enterprise memory, LongMemEval
Zep implements Graphiti, a temporally-aware knowledge graph where each fact is timestamped and can evolve over time. The system integrates both unstructured conversational data and structured business data in a single relational graph — enabling temporal queries like 'what happened before session 3?' or 'how did the learner's understanding of this concept evolve?'.
Keywords: full-duplex, multilingual EN/FR, EU sovereignty, GDPR by design
Gradium is a production-optimized Kyutai spin-off: architecture derived from Moshi with streamed neural Mimi codec and two parallel audio streams. Multilingual EN/FR focus from the start, EU hosting, GDPR by design compliance — 'The voice infrastructure company for Europe'. Limited API access in May 2026, French WER not yet officially published.
Keywords: unified text→avatar pipeline, dual-branch diffusion transformer, zero-shot 25 FPS
OmniTalker is the first unified end-to-end framework that simultaneously generates speech and talking head video from text and a reference video, in zero-shot real-time inference (25 FPS). Dual-branch diffusion transformer architecture: audio branch synthesizes mel-spectrograms from text, visual branch predicts poses and facial dynamics. Audio-visual fusion module ensures synchronization and stylistic coherence.
Keywords: epistemic trust, learner identity, perceived competence/benevolence/integrity
Integrative theory-building review synthesizing literature on epistemic trust, learner identity formation, and organizational readiness in avatar-mediated instruction contexts in higher education. Develops a framework articulating how AI avatars influence learner engagement and reconfigure digital identity formation. Epistemic trust — belief that the avatar is a reliable source — is identified as the central mediator.
Keywords: photorealistic talking head, nuanced expressions, real-time
VASA-1 demonstrates that photorealistic talking faces can be generated at 40 FPS in real-time from a single image, with nuanced emotional expressions going well beyond simple lip-sync. Not yet commercialized (risk of incomplete publication), this work establishes a quality benchmark reference for expressive avatars.
Keywords: zero-shot TTS, style encoder, individual prosody, spontaneous speech
PerTTS proposes a personalized and controllable zero-shot voice synthesis system: a speech style encoder captures the individual prosodic fingerprint (rhythm, intonation, spontaneous patterns) from a few audio samples, and a local prosody encoder controls phrase-by-phrase variations. The approach requires no full fine-tuning — a few minutes of audio suffice.
Keywords: disentangled facial latent space, 512×512 @ 40 FPS generation, negligible latency
VASA-1 generates talking avatars from a single static image and an audio clip. The disentangled facial latent space treats lip movements, expressions, gaze, and head movements as independent variables. The diffusion model operates in this latent space, enabling online generation of 512×512 videos at 40 FPS. Microsoft did not release this model publicly, acknowledging the risk of misuse.
Keywords: agentic memory, temporal knowledge graph, conversational agents
This paper proposes an agentic memory architecture based on temporal knowledge graphs, where each graph node represents a memorized fact with its timestamp, relations to other facts, and obsolescence probability. The approach enables temporal queries on the evolution of a user's cognitive state — crucial for long-term pedagogical tracking scenarios.
Keywords: audio-driven diffusion, long-form streaming, GPU cost, open source
LiveAvatar jointly designs algorithm and system for continuous audio-driven sequences. The repository documents substantial acceleration, but also dependence on heavy GPU infrastructure and TTS integration still listed as a next step.
Keywords: RAG, vector DB memory, conversational LLMs, survey
Systematic review of retrieval-augmented memory architectures (RAG) for conversational LLMs. The article synthesizes vector database approaches, compares chunking and indexing strategies, and evaluates trade-offs between retrieval latency and contextual accuracy.
Keywords: RAG, long-term memory, agentic agents, reinforcement learning
This paper documents the transition from static RAG architectures to agentic long-term memory systems, where the agent learns to manage its own memory via reinforcement learning. The approach enables active selection of memories to retain or forget, rather than simple mechanical compression.
Keywords: complete digital human, 3D avatar, expressive speech, grounded dialogue
Hi-Reco proposes a rare integrated approach: a complete digital human combining 3D avatar, expressive speech, and grounded dialogue in a unified architecture. Unlike modular pipelines (separate TTS + avatar), the end-to-end approach guarantees coherence between vocal and facial expression.
Keywords: talking head synthesis, comprehensive survey, real-time / expressiveness / quality trilemma
Comprehensive review of talking head synthesis techniques published in ACM Computing Surveys (2025). The article formally documents the fundamental trilemma: no current solution simultaneously satisfies all three constraints — real-time (<500ms), rich emotional expressiveness, and photorealistic quality. This trilemma structures the research space.
Keywords: TTS, complex style control, benchmark, evaluation 11Labs / Deepgram / OpenAI
EmergentTTS-Eval is a NeurIPS 2025 benchmark for evaluating complex style control in voice synthesis. It evaluates 11 expressive dimensions (emotion, intensity, rhythm, character) on major market TTS systems (ElevenLabs, Deepgram, OpenAI 4o-mini-TTS) and establishes reproducible comparative metrics.
Keywords: multimodal speech model, contextual prosody, natural backchannels, RVQ
Sesame CSM is a multimodal conversational speech model designed to 'cross the uncanny valley of conversational voice': it generates RVQ audio codes from text and audio inputs, with contextual prosody, natural backchannels ('mm-hm', 'yes') and human turn-taking behavior. Not optimized for production streaming but a high-quality research reference.
Keywords: tandem architecture, agentic memory, knowledge graph, real-time RAG
The KAME Tandem architecture combines a dynamic knowledge graph with real-time agentic memory. Its 'tandem' approach couples two complementary components: a fast retrieval module for immediate facts, and a deep reasoning module for complex semantic relationships. This is the Slow component of the MoshiRAG architecture in GamiWays taxonomy.
Keywords: Timestep-forcing Pipeline Parallelism, Rolling Sink Frame, identity drift, 14B parameters
Live Avatar resolves the video diffusion generation bottleneck for real-time applications. Timestep-forcing Pipeline Parallelism (TPP) pipelines denoising steps across multiple GPUs. The Rolling Sink Frame Mechanism (RSFM) recalibrates appearance from a cached reference image, eliminating identity drift on long sequences. 20 FPS end-to-end on 5 H800 GPUs with 14B parameters.
Keywords: Anchor-Heavy Identity Sinks, 20x inference cost reduction, multi-turn, outperforms Sora2/Veo3
LiveTalk combines a video diffusion model conditioned on text, image, and audio with audio language models. Anchor-Heavy Identity Sinks technique for long-duration inference. The distilled model achieves comparable quality to bidirectional baselines with 20x less inference cost. In multi-turn evaluation, LiveTalk outperforms Sora2 and Veo3 in video coherence, bringing latency from 1-2 minutes to real time.
Keywords: comprehensive survey, GANs→diffusion transformers, HDTF/VFHQ/LRS3 benchmarks, ethnic diversity
Systematic review covering the evolution of talking head generation techniques from GANs (Wav2Lip, SadTalker) to recent diffusion transformers. Maps standard benchmarks (HDTF, VFHQ, LRS3), evaluation metrics (FID, FVD, SyncNet, PSNR), and architectural trade-offs. Identifies four priority improvement axes: identity on long sequences, expressiveness beyond lips, ethnic and cultural diversity, inference speed on constrained hardware.
Keywords: HRV/EEG, sympathetic stress, presence-anonymity trilemma, low/medium/high visual fidelity
Experimental study measuring physiological (HRV, EEG) and psychological (perceived stress) responses of participants interacting with low, medium, and high visual fidelity avatars. High-fidelity avatars generate a significantly higher LF/HF ratio (sympathetic stress) than low-fidelity avatars, particularly during sensitive conversations. Confirms the presence-anonymity trilemma.
Keywords: parasocial interaction, digital attachment, emotional dependency, ethical design
Systematic review of two decades of research proposing a theoretical model of the evolution of human-AI emotional relationships: from parasocial interaction (unilateral emotional investment) to digital attachment. Identified determinants: AI personalization, memory continuity, perceived emotional reactivity, anthropomorphization. Dependency risks and ethical design criteria proposed.
Keywords: pseudo-intimacy, affective AI, anthropomorphization bias, transparency, informed consent
Explores the risk of pseudo-intimacy with affective AIs: emotional AIs create relationships that mimic intimacy without the conditions of reciprocity, shared vulnerability, and common history that found authentic human relationships. Analyzes cognitive mechanisms (anthropomorphization bias, affective reinforcement by simulated reactivity) and proposes an affective design ethics framework based on transparency, deliberate expressiveness limitation, and informed consent.
Keywords: educational avatars, metaverse, generative NLP, data security, algorithmic bias
Systematic review evaluating avatar contributions to learning in immersive virtual environments. Documents that avatars enable individualized experiences and improved collaborative activities, and that advances in NLP and generative models have significantly improved their capabilities. Identifies major obstacles: data security, ethical concerns (algorithmic bias, privacy), and limited infrastructure.
Keywords: ethnic bias, DH-FaceVid-1K dataset, multilingual generalization, algorithmic fairness, cultural representation
This ACM Grand Challenge directly targets the major gap in talking head generation models: their poor generalization on non-European and non-American ethnicities. The DH-FaceVid-1K dataset is introduced to address this gap. Participants must develop models integrating diffusion and GAN techniques to improve lip synchronization, identity coherence, and inference speed on multilingual and multi-ethnic data. This challenge reveals that standard benchmarks (HDTF, VFHQ) are massively biased toward Caucasian English-speaking faces.
Keywords: long-term memory benchmark, LLM assistants, systematic evaluation
This benchmark establishes a systematic evaluation framework for long-term memory capabilities of LLM assistants, covering fact retention, temporal coherence, and progressive personalization. It documents current limits of naive approaches (full context, truncation) and opens the path toward genuinely personalized long-term assistants.
Keywords: context compression, LLM inference, long sessions, 128K tokens
This research proposes stateful context compression enabling a 128K token attention window without perceptible quality degradation. Unlike brutal truncation, the stateful approach preserves a compressed representation of history — particularly applicable to long conversational sessions (1h+).
Keywords: digital twin ownership, data autonomy, renewed social contract, post-mortem rights
Argues that natural persons must be recognized as moral and legal owners of their AI digital twins, as these entities are intimate extensions of the individual built from their personal data. Current legal frameworks, which privilege technological infrastructure over data autonomy, are criticized as perpetuating systemic inadequacies. Proposes a renewed social contract centered on individual dominion and consent.
Keywords: vocal uncanny valley, neural TTS vs human voice, trust, speaker/listener gender
This ACM CHI study demonstrates that a new dimension of the uncanny valley emerges in the vocal domain: a virtual human using neural TTS is perceived as significantly less trustworthy than a virtual human using a real human voice. The effect is modulated by speaker and listener gender. The study extends Mori's (1970) classic concept from visual morphology to acoustic fidelity.
Keywords: Audio-Pose Prior Refocusing, rhythm-gesture coupling, bidirectional cross-attention
ApoAvatar introduces an Audio-Pose Prior Refocusing mechanism that links speech style to movement dynamics: strong accents amplify gesture amplitude, silent passages suppress superfluous movements. A frame-wise audio-video interaction module updates audio features with current visual context via bidirectional cross-attention. Clear gains on lip synchronization, gesture expressiveness, and overall naturalness.
Keywords: speech-synchronized hand gestures, audio-driven poses, perceived authenticity
EMO2 extends audio-driven avatar generation beyond the face to include speech-synchronized hand gestures. Two-step process: hand pose generation directly from audio (strong audio-gesture correlation), then diffusion model synthesizing video frames integrating these poses. Outperforms CyberHost and Vlogger on visual quality and synchronization accuracy.
Keywords: EU AI Act Art. 50(2), mandatory labeling, €35M penalties, human rights, prohibited emotional recognition
Analyzes how the EU AI Act (adopted 2024, full enforcement August 2026) establishes a rights-centered framework for regulating synthetic media. Article 50(2) requires that generative outputs including deepfakes be clearly and visibly labeled. Penalties: up to €35M or 7% of global revenue. Emotional recognition in educational settings is explicitly prohibited by Article 5 since February 2025.
Keywords: AI inference carbon footprint, 0.24 Wh/Gemini query, video generation 10–100× more costly, IDC 23 TWh
MIT Technology Review's analysis evaluates AI energy footprint at sector scale: IDC estimates AI data centers consumed 23 TWh in 2022. Google publishes a 2025 methodology for measuring AI inference environmental impact: a median Gemini query consumes 0.24 Wh and emits 0.03 g CO₂. Large open-source models (70B parameters) consume up to 1.7 Wh per query. Video generation is significantly more costly than text generation — likely 10 to 100 times more energy-intensive for real-time avatar streaming.
Browse the detailed domain-by-domain classification
This view keeps the corpus’ historical organization. Use the filters above first to isolate a decision, then return here to read references by their original academic domain.
7 publications covering memory stratification, context compression, RAG architectures, and temporal knowledge graphs. Directly applicable to Core Engine hypothesis H2.
Keywords: STM / LTM / Knowledge layers, personalized AI agents
MemoryOS proposes a memory 'operating system' for personalized AI agents, organized in three distinct layers: short-term working memory, long-term episodic memory, and a structured knowledge base. The accuracy gains (+37.6% vs baseline) experimentally validate that memory stratification outperforms a flat context window.
Keywords: persistent memory, token reduction, conversational agents
Mem0 is an open-source persistent structured memory library for AI agents. Its mechanism relies on automatic extraction of salient facts (preferences, events, context) at each conversation turn, followed by intelligent deduplication and targeted retrieval — only relevant memories are injected into the prompt at each exchange.
Keywords: long-term memory benchmark, LLM assistants, systematic evaluation
This benchmark establishes a systematic evaluation framework for long-term memory capabilities of LLM assistants, covering fact retention, temporal coherence, and progressive personalization. It documents current limits of naive approaches (full context, truncation) and opens the path toward genuinely personalized long-term assistants.
Keywords: context compression, LLM inference, long sessions, 128K tokens
This research proposes stateful context compression enabling a 128K token attention window without perceptible quality degradation. Unlike brutal truncation, the stateful approach preserves a compressed representation of history — particularly applicable to long conversational sessions (1h+).
Keywords: RAG, vector DB memory, conversational LLMs, survey
Systematic review of retrieval-augmented memory architectures (RAG) for conversational LLMs. The article synthesizes vector database approaches, compares chunking and indexing strategies, and evaluates trade-offs between retrieval latency and contextual accuracy.
Keywords: RAG, long-term memory, agentic agents, reinforcement learning
This paper documents the transition from static RAG architectures to agentic long-term memory systems, where the agent learns to manage its own memory via reinforcement learning. The approach enables active selection of memories to retain or forget, rather than simple mechanical compression.
Keywords: temporal knowledge graph, enterprise memory, LongMemEval
Zep implements Graphiti, a temporally-aware knowledge graph where each fact is timestamped and can evolve over time. The system integrates both unstructured conversational data and structured business data in a single relational graph — enabling temporal queries like 'what happened before session 3?' or 'how did the learner's understanding of this concept evolve?'.
8 publications covering photorealistic avatars, personalized voice synthesis, TTS benchmarks, and multimodal speech models. Directly applicable to R&D axes 1, 2a, and 3.
Keywords: photorealistic talking head, nuanced expressions, real-time
VASA-1 demonstrates that photorealistic talking faces can be generated at 40 FPS in real-time from a single image, with nuanced emotional expressions going well beyond simple lip-sync. Not yet commercialized (risk of incomplete publication), this work establishes a quality benchmark reference for expressive avatars.
Keywords: audio-avatar LLM, expressive facial movements, LoRA architecture
A²-LLM is an end-to-end audio-avatar LLM that generates emotionally rich facial movements beyond lip-sync, by directly coupling text generation and facial synthesis in a single model. The 8B + 0.16B LoRA architecture enables fine adaptation to different expressive styles without retraining the entire model.
Keywords: complete digital human, 3D avatar, expressive speech, grounded dialogue
Hi-Reco proposes a rare integrated approach: a complete digital human combining 3D avatar, expressive speech, and grounded dialogue in a unified architecture. Unlike modular pipelines (separate TTS + avatar), the end-to-end approach guarantees coherence between vocal and facial expression.
Keywords: talking head synthesis, comprehensive survey, real-time / expressiveness / quality trilemma
Comprehensive review of talking head synthesis techniques published in ACM Computing Surveys (2025). The article formally documents the fundamental trilemma: no current solution simultaneously satisfies all three constraints — real-time (<500ms), rich emotional expressiveness, and photorealistic quality. This trilemma structures the research space.
Keywords: streaming diffusion, Local-Future Sliding-Window Denoising, long-form real-time
AvatarForcing is the most directly applicable academic paper to R&D Axis 1 (latency reduction). Its 'Local-Future Sliding-Window Denoising' technique enables talking avatar video generation in 1-step streaming diffusion — drastically reducing denoising iterations (from 20–50 steps to 1 step), dividing generation time by 60–90%.
Keywords: TTS, complex style control, benchmark, evaluation 11Labs / Deepgram / OpenAI
EmergentTTS-Eval is a NeurIPS 2025 benchmark for evaluating complex style control in voice synthesis. It evaluates 11 expressive dimensions (emotion, intensity, rhythm, character) on major market TTS systems (ElevenLabs, Deepgram, OpenAI 4o-mini-TTS) and establishes reproducible comparative metrics.
Keywords: zero-shot TTS, style encoder, individual prosody, spontaneous speech
PerTTS proposes a personalized and controllable zero-shot voice synthesis system: a speech style encoder captures the individual prosodic fingerprint (rhythm, intonation, spontaneous patterns) from a few audio samples, and a local prosody encoder controls phrase-by-phrase variations. The approach requires no full fine-tuning — a few minutes of audio suffice.
Keywords: multimodal speech model, contextual prosody, natural backchannels, RVQ
Sesame CSM is a multimodal conversational speech model designed to 'cross the uncanny valley of conversational voice': it generates RVQ audio codes from text and audio inputs, with contextual prosody, natural backchannels ('mm-hm', 'yes') and human turn-taking behavior. Not optimized for production streaming but a high-quality research reference.
6 recent works (2025–2026) documenting full-duplex interaction models, agentic memory architectures, and European sovereignty solutions. These works inform GamiWays' Phase B/C architectural choices.
Architecture: MoE 276B/12B active, native dMel audio tokenization, dual Fast/Slow model
TML Interaction-Small is a MoE model with 276B total parameters but only 12B active per inference. It integrates native dMel audio tokenization and a dual architecture: the Fast Model (200ms chunks) handles real-time conversation while the Slow Model (asynchronous) handles tool calls, RAG, and deep reasoning. Visual proactivity and absolute temporal awareness. Known limitation: coherence degrading after ~30–45 minutes.
Keywords: full-duplex, asynchronous RAG, Moshi, enriched memory
MoshiRAG is an extension of Moshi (Kyutai) integrating asynchronous RAG and enriched memory in a full-duplex model. The target architecture is exactly GamiWays' Phase C: full-duplex + RAG/retrieval/reasoning in parallel, without blocking the audio stream. The GamiWays Core Engine is architecturally this orchestration layer.
Keywords: tandem architecture, agentic memory, knowledge graph, real-time RAG
The KAME Tandem architecture combines a dynamic knowledge graph with real-time agentic memory. Its 'tandem' approach couples two complementary components: a fast retrieval module for immediate facts, and a deep reasoning module for complex semantic relationships. This is the Slow component of the MoshiRAG architecture in GamiWays taxonomy.
Keywords: duplex cascade, conversational pipeline, reduced latency, streaming
DuplexCascade proposes a duplex cascade conversational pipeline architecture, reducing perceived latency by starting each step (TTS, avatar generation) from the first tokens of the previous step, without waiting for completion. This is the academic formalization of the 'streaming pipeline' approach GamiWays targets for Phase A+ (estimated gain: −40% total latency).
Keywords: agentic memory, temporal knowledge graph, conversational agents
This paper proposes an agentic memory architecture based on temporal knowledge graphs, where each graph node represents a memorized fact with its timestamp, relations to other facts, and obsolescence probability. The approach enables temporal queries on the evolution of a user's cognitive state — crucial for long-term pedagogical tracking scenarios.
Keywords: full-duplex, multilingual EN/FR, EU sovereignty, GDPR by design
Gradium is a production-optimized Kyutai spin-off: architecture derived from Moshi with streamed neural Mimi codec and two parallel audio streams. Multilingual EN/FR focus from the start, EU hosting, GDPR by design compliance — 'The voice infrastructure company for Europe'. Limited API access in May 2026, French WER not yet officially published.
20 references covering 6 dimensions: technical foundations (VASA-1, RETA, AvatarForcing), psychology (uncanny valley, attachment, epistemic trust), social and pedagogical implications, law and regulation (EU AI Act, biometric data, avatar ownership), environmental impact and sustainability, and cross-cutting critical perspectives (ethnic bias).
Keywords: disentangled facial latent space, 512×512 @ 40 FPS generation, negligible latency
VASA-1 generates talking avatars from a single static image and an audio clip. The disentangled facial latent space treats lip movements, expressions, gaze, and head movements as independent variables. The diffusion model operates in this latent space, enabling online generation of 512×512 videos at 40 FPS. Microsoft did not release this model publicly, acknowledging the risk of misuse.
Keywords: real-time / sync / emotion trilemma, 3DMM, single-pass GAN, 55+ FPS
RETA is the first framework to simultaneously resolve the fundamental trilemma: real-time performance, lip synchronization accuracy, and emotional fidelity. It disentangles the audio signal into two representations: a 3DMM geometry for lip synchronization, and a dynamic emotional embedding learned without labels via cross-modal distillation. A single-pass GAN generator integrates these representations hierarchically. 55+ FPS with new SOTA on all metrics.
Keywords: Timestep-forcing Pipeline Parallelism, Rolling Sink Frame, identity drift, 14B parameters
Live Avatar resolves the video diffusion generation bottleneck for real-time applications. Timestep-forcing Pipeline Parallelism (TPP) pipelines denoising steps across multiple GPUs. The Rolling Sink Frame Mechanism (RSFM) recalibrates appearance from a cached reference image, eliminating identity drift on long sequences. 20 FPS end-to-end on 5 H800 GPUs with 14B parameters.
Keywords: Local-Future Sliding-Window, 1.3B parameters, 34 ms/frame, 2-step distillation
AvatarForcing resolves the exposure bias of autoregressive models via a sliding window with heterogeneous denoisings (Local-Future Sliding Window). A dual temporal anchoring — style anchor (RoPE re-indexing) and temporal anchor (recent clean block reuse) — stabilizes the stream over unlimited durations. The student model (1.3B parameters) achieves 34 ms/frame while maintaining strong visual quality on a 400 long-duration video benchmark.
Keywords: diffusion forcing, non-verbal reactivity, synthetic DPO, 500ms latency, 6.8x speedup
Avatar Forcing models user-avatar interactions through diffusion forcing, allowing the avatar to process the user's multimodal inputs (audio, head movement, expressions) in real time with approximately 500ms latency. A direct preference optimization method with synthetic losing samples enables expressive learning without additional annotations. 6.8x speedup vs baseline.
Keywords: unified text→avatar pipeline, dual-branch diffusion transformer, zero-shot 25 FPS
OmniTalker is the first unified end-to-end framework that simultaneously generates speech and talking head video from text and a reference video, in zero-shot real-time inference (25 FPS). Dual-branch diffusion transformer architecture: audio branch synthesizes mel-spectrograms from text, visual branch predicts poses and facial dynamics. Audio-visual fusion module ensures synchronization and stylistic coherence.
Keywords: speech-synchronized hand gestures, audio-driven poses, perceived authenticity
EMO2 extends audio-driven avatar generation beyond the face to include speech-synchronized hand gestures. Two-step process: hand pose generation directly from audio (strong audio-gesture correlation), then diffusion model synthesizing video frames integrating these poses. Outperforms CyberHost and Vlogger on visual quality and synchronization accuracy.
Keywords: Audio-Pose Prior Refocusing, rhythm-gesture coupling, bidirectional cross-attention
ApoAvatar introduces an Audio-Pose Prior Refocusing mechanism that links speech style to movement dynamics: strong accents amplify gesture amplitude, silent passages suppress superfluous movements. A frame-wise audio-video interaction module updates audio features with current visual context via bidirectional cross-attention. Clear gains on lip synchronization, gesture expressiveness, and overall naturalness.
Keywords: Anchor-Heavy Identity Sinks, 20x inference cost reduction, multi-turn, outperforms Sora2/Veo3
LiveTalk combines a video diffusion model conditioned on text, image, and audio with audio language models. Anchor-Heavy Identity Sinks technique for long-duration inference. The distilled model achieves comparable quality to bidirectional baselines with 20x less inference cost. In multi-turn evaluation, LiveTalk outperforms Sora2 and Veo3 in video coherence, bringing latency from 1-2 minutes to real time.
Keywords: comprehensive survey, GANs→diffusion transformers, HDTF/VFHQ/LRS3 benchmarks, ethnic diversity
Systematic review covering the evolution of talking head generation techniques from GANs (Wav2Lip, SadTalker) to recent diffusion transformers. Maps standard benchmarks (HDTF, VFHQ, LRS3), evaluation metrics (FID, FVD, SyncNet, PSNR), and architectural trade-offs. Identifies four priority improvement axes: identity on long sequences, expressiveness beyond lips, ethnic and cultural diversity, inference speed on constrained hardware.
Keywords: vocal uncanny valley, neural TTS vs human voice, trust, speaker/listener gender
This ACM CHI study demonstrates that a new dimension of the uncanny valley emerges in the vocal domain: a virtual human using neural TTS is perceived as significantly less trustworthy than a virtual human using a real human voice. The effect is modulated by speaker and listener gender. The study extends Mori's (1970) classic concept from visual morphology to acoustic fidelity.
Keywords: HRV/EEG, sympathetic stress, presence-anonymity trilemma, low/medium/high visual fidelity
Experimental study measuring physiological (HRV, EEG) and psychological (perceived stress) responses of participants interacting with low, medium, and high visual fidelity avatars. High-fidelity avatars generate a significantly higher LF/HF ratio (sympathetic stress) than low-fidelity avatars, particularly during sensitive conversations. Confirms the presence-anonymity trilemma.
Keywords: parasocial interaction, digital attachment, emotional dependency, ethical design
Systematic review of two decades of research proposing a theoretical model of the evolution of human-AI emotional relationships: from parasocial interaction (unilateral emotional investment) to digital attachment. Identified determinants: AI personalization, memory continuity, perceived emotional reactivity, anthropomorphization. Dependency risks and ethical design criteria proposed.
Keywords: pseudo-intimacy, affective AI, anthropomorphization bias, transparency, informed consent
Explores the risk of pseudo-intimacy with affective AIs: emotional AIs create relationships that mimic intimacy without the conditions of reciprocity, shared vulnerability, and common history that found authentic human relationships. Analyzes cognitive mechanisms (anthropomorphization bias, affective reinforcement by simulated reactivity) and proposes an affective design ethics framework based on transparency, deliberate expressiveness limitation, and informed consent.
Keywords: educational avatars, metaverse, generative NLP, data security, algorithmic bias
Systematic review evaluating avatar contributions to learning in immersive virtual environments. Documents that avatars enable individualized experiences and improved collaborative activities, and that advances in NLP and generative models have significantly improved their capabilities. Identifies major obstacles: data security, ethical concerns (algorithmic bias, privacy), and limited infrastructure.
Keywords: epistemic trust, learner identity, perceived competence/benevolence/integrity
Integrative theory-building review synthesizing literature on epistemic trust, learner identity formation, and organizational readiness in avatar-mediated instruction contexts in higher education. Develops a framework articulating how AI avatars influence learner engagement and reconfigure digital identity formation. Epistemic trust — belief that the avatar is a reliable source — is identified as the central mediator.
Keywords: EU AI Act Art. 50(2), mandatory labeling, €35M penalties, human rights, prohibited emotional recognition
Analyzes how the EU AI Act (adopted 2024, full enforcement August 2026) establishes a rights-centered framework for regulating synthetic media. Article 50(2) requires that generative outputs including deepfakes be clearly and visibly labeled. Penalties: up to €35M or 7% of global revenue. Emotional recognition in educational settings is explicitly prohibited by Article 5 since February 2025.
Keywords: digital twin ownership, data autonomy, renewed social contract, post-mortem rights
Argues that natural persons must be recognized as moral and legal owners of their AI digital twins, as these entities are intimate extensions of the individual built from their personal data. Current legal frameworks, which privilege technological infrastructure over data autonomy, are criticized as perpetuating systemic inadequacies. Proposes a renewed social contract centered on individual dominion and consent.
Keywords: AI inference carbon footprint, 0.24 Wh/Gemini query, video generation 10–100× more costly, IDC 23 TWh
MIT Technology Review's analysis evaluates AI energy footprint at sector scale: IDC estimates AI data centers consumed 23 TWh in 2022. Google publishes a 2025 methodology for measuring AI inference environmental impact: a median Gemini query consumes 0.24 Wh and emits 0.03 g CO₂. Large open-source models (70B parameters) consume up to 1.7 Wh per query. Video generation is significantly more costly than text generation — likely 10 to 100 times more energy-intensive for real-time avatar streaming.
Keywords: ethnic bias, DH-FaceVid-1K dataset, multilingual generalization, algorithmic fairness, cultural representation
This ACM Grand Challenge directly targets the major gap in talking head generation models: their poor generalization on non-European and non-American ethnicities. The DH-FaceVid-1K dataset is introduced to address this gap. Participants must develop models integrating diffusion and GAN techniques to improve lip synchronization, identity coherence, and inference speed on multilingual and multi-ethnic data. This challenge reveals that standard benchmarks (HDTF, VFHQ) are massively biased toward Caucasian English-speaking faces.
Summary — Maturity levels by domain
Source: gamiways-recherches-2026.md — May 2026
| Domain | Academic maturity | Commercial availability | GamiWays gap |
|---|---|---|---|
| Conversational memory | High (Mem0, MemoryOS) | Partial (Mem0 API) | 3-layer architecture + avatar-specific SLM distillation |
| Multi-session avatar integration | High (VASA-1, A²-LLM) | Partial (HeyGen, Tavus) | Latency <500ms + full body |
| Personalized expressive TTS | High (PerTTS, EmergentTTS) | Good (ElevenLabs, Cartesia) | Individual prosodic fingerprint |
| Conversational orchestration | Emerging (MoshiRAG, KAME) | Low | Configurable freedom spectrum + agentic Game Master |
Open Research Questions
These unresolved questions structure the GamiWays R&D program 2026–2028.