HeyGen LiveAvatar + v3 Video API
Avatar temps réel (Custom/LITE ou FULL) + pipeline v3 asynchrone (Video Agent, Avatar V, Cinematic)
TTFR Latency
~300ms
real-time
Cost / minute
from $0.095/min
depending on plan
Visual Quality
9/10
estimated score
Protocols
WebRTC, REST, LiveKit, WebSocket
Two separate developer surfaces
HeyGen expose deux surfaces distinctes avec des clés incompatibles. LiveAvatar (api.liveavatar.com) gère les sessions WebRTC temps réel, en FULL ou Custom/LITE. HeyGen v3 (api.heygen.com) gère la vidéo asynchrone, la traduction, le TTS et les workflows agentiques. Où est Ava ? utilise le mode Custom/LITE pour conserver sa pile vocale et v3 pour des segments cinématiques pré-rendus.
Scene & Staging Control via API
Direct comparison between real-time (LiveAvatar) and async video generation (HeyGen v3). This is the key question for Où est Ava ? and Dilemme Plastique.
| Aspect | Real-time (LiveAvatar) | Async (HeyGen v3) |
|---|---|---|
| Background / décor | ✗ Fixed by avatar training video — no API parameter | ✓ background (color/image), remove_background on Direct Video |
| Camera / cadrage | ✗ Not controllable | ✓ Cinematic Avatar prompt (handheld, documentary, etc.) — async only |
| Body motion / gestures | ✗ Emergent from model — not programmable | ✓ motion_prompt (Avatar IV) · Cinematic prompt · Video Agent prompt |
| Facial expressiveness | ✗ Automatic from training | ✓ expressiveness: high|medium|low (Avatar IV photo/image only) |
| Multi-avatar scenes | ✗ Single avatar per session | ✓ Cinematic Avatar: 1–3 looks in one shot |
| Personality / dialogue | ✓ Context object (FULL) or external LLM (LITE) | ✓ script (Direct) or agent-written (Video Agent) |
| Reference media / brand assets | ✗ | ✓ files[] in Video Agent · references[] in Cinematic · style_id |
| Real-time lip-sync to custom TTS | ✓ Custom/LITE — vidéo streamée au-dessus de votre pile vocale | ✗ Async render only (3–8s+ per clip in benchmarks) |
HeyGen v3 + LiveAvatar API Products
LiveAvatar Custom (LITE)
Low controlTTFF <300 ms annoncé par le fournisseur — mesure de dialogue Où est Ava ? à réaliserPOST /v1/sessions/token + /sessions/start
Meilleur mode à tester pour GamiWays : conserver STT/LLM/RAG/TTS et l’orchestration, LiveAvatar ne rendant que la vidéo synchronisée. LiveKit est fourni par défaut ; une infrastructure LiveKit ou Agora propre est possible.
- →avatar_id — apparence visuelle, catalogue ou avatar custom (fixé par le matériau d’entraînement)
- →quality — very_high (1080p) | high (720p) | medium (480p) | low (360p)
- →encoding — H264 recommandé (VP8 déprécié)
- →activity_idle_timeout / disable_idle_timeout
- →No background, camera, gesture, or expression API — apporter la direction vocale via votre TTS
LiveAvatar FULL
Low control~400–600ms TTFRPOST /v1/sessions/token (mode: FULL)
End-to-end managed pipeline (VAD + STT + LLM + TTS + video). Faster to ship, less control over conversation stack.
- →avatar_id, voice_id, context_id
- →avatar_persona.language, voice_settings (speed, stability)
- →interactivity_type — conversational | push_to_talk
- →Context: system prompt, opening text, RAG links, guardrails
- →No scene/staging parameters beyond avatar choice
Direct Video (Avatar IV/V)
Structured controlAsync — seconds to minutesPOST /v3/videos (type: avatar)
Explicit script + avatar + voice. Highest structured control for talking-head async content. Avatar V = better motion/lip-sync; Avatar IV adds motion_prompt + expressiveness.
- →avatar_id, voice_id, script
- →background — solid color or image URL
- →remove_background — boolean (requires matting-enabled twin)
- →motion_prompt — natural language body motion (Avatar IV only)
- →expressiveness — high | medium | low (Avatar IV only)
- →aspect_ratio — auto | 16:9 | 9:16 | 4:5 | 5:4 | 1:1
- →resolution — 720p | 1080p | 4k
- →voice_settings — speed, pitch, locale
- →engine — avatar_iv (default) | avatar_v
Video Agent
Prompt-onlyAsync — agent plans then rendersPOST /v3/video-agents
Flagship v3 workflow: one prompt → script + avatar + scenes. Good for rapid prototyping of educational or marketing segments, less predictable output.
- →prompt (1–10,000 chars) — tone, audience, pacing, visual style
- →style_id — curated visual templates (GET /v3/video-agents/styles)
- →orientation — landscape | portrait
- →avatar_id, voice_id — optional overrides
- →files — up to 20 reference assets (slides, images, PDFs)
- →mode: chat — interactive storyboard before render
- →callback_url — webhook instead of polling
Cinematic Avatar
High staging controlAsync — 4–15s clipsPOST /v3/videos (type: cinematic_avatar)
Strongest scene/staging control in HeyGen API — but async, no spoken voice. Prompt-driven camera, setting, mood, and motion via Seedance pipeline.
- →prompt — scene, action, camera, mood (1–10,000 chars)
- →avatar_id — array of 1–3 looks (multi-avatar shots)
- →references — up to 3 videos + 9 images for style/motion steering
- →duration — 4–15s | auto_duration
- →enhance_prompt — auto-expand short prompts
- →aspect_ratio — 16:9 | 9:16 | 1:1
- →Flat $7/video pricing
HyperFrames
High staging controlAsyncPOST /v3/hyperframes/renders
HTML/CSS/JS → motion graphics video. Useful for Dilemme Plastique data visualizations or Où est Ava ? UI overlays — not a talking avatar.
- →HTML composition input
- →resolution — 1080p | 4k
- →aspect_ratio — 16:9 | 9:16 | 1:1
Video Translation v3
Low controlAsyncPOST /v3/video-translations
Translate existing video into 30+ languages with voice cloning + lip-sync. Relevant for localizing Dilemme Plastique or Où est Ava ? pre-rendered segments.
- →mode — speed | precision
- →brand_glossary_id — terminology consistency
- →Proofread sessions before final render (precision mode)
GamiWays Project Applications
Où est Ava ? (Storygami) — Où est Ava ?
8/10LiveAvatar Custom/LITE + v3 Cinematic Avatar (B-roll)
Architecture hybride : le mode Custom/LITE de LiveAvatar préserve la pile ElevenLabs + LLM et le RAG d’Où est Ava ? ; le fournisseur annonce un TTFF inférieur à 300 ms et une concurrence illimitée dès Starter, à mesurer en condition réelle. Des inserts cinématiques pré-rendus restent nécessaires pour les transitions, flashbacks ou changements de décor, car la mise en scène temps réel demeure verrouillée par l’avatar entraîné.
- No real-time scene changes during conversation — staging is locked to training video
- Cinematic clips are 4–15s, voiceless — must be edited into the experience
- US hosting only — GDPR review for Swiss/EU deployment
- Interactive Avatar API sunset March 31, 2026 — migrate to LiveAvatar
Le Dilemme Plastique (Edugami)
7/10LiveAvatar Custom/LITE + v3 Direct Video + Translation
Le flux pédagogique voice-first correspond au mode Custom/LITE : conserver le STT, LLM/RAG et TTS français choisis, puis ajouter la couche de lip-sync LiveAvatar. La concurrence illimitée annoncée rend le test de cohortes plus crédible, sans lever la nécessité de mesurer qualité, coûts et durées de session. La vidéo asynchrone reste adaptée aux explications scénarisées avec diapositives.
- Scientific credibility requires stable avatar — train once, reuse avatar_id
- expressiveness/motion_prompt only on Avatar IV — test quality vs Avatar V
- HyperFrames candidate for data/chart animations — separate from avatar layer
- Async segments incompatible with sub-2s turn latency target for live Q&A
Emotional Direction & Staging
LiveAvatar Custom/LITE lets you select a custom avatar and bring directed voice, but provides no programmable streaming gesture or expression; advanced control belongs to async v3 products.
Limited controlParameters and mechanism
Async Avatar IV: `motion_prompt`, `expressiveness`; LiveAvatar: `avatar_id` + external audio stream in Custom/LITE, with no per-turn gesture/emotion API.
Limit to keep visible
Do not confuse async control with streaming dialogue: staged segments can complement, not direct every live turn. Glossary and professional cloning have their own access and consent terms.
Avatar Customisation
RAG / Knowledge Base
FULL mode : contexte et connaissances fournis par LiveAvatar. Custom/LITE : votre propre RAG, vos garde-fous et votre orchestration ; LiveAvatar rend la vidéo uniquement.
Behavior & Personality
FULL mode : voice agent, contexte, garde-fous et conversation/push-to-talk. Custom/LITE : contrôle complet via votre LLM et orchestration ; la direction émotionnelle doit venir du TTS/agent retenu, LiveAvatar n’ajoute pas de couche cognitive.
Body Language & Gestures
Real-time: not programmable — fixed by training video, model generates natural nods/eye contact. Async: motion_prompt (Avatar IV) or Cinematic prompt for composed shots.
Facial Expressions
Real-time: automatic from training. Async Avatar IV: expressiveness high/medium/low on photo/image avatars only. No per-expression API.
Voice & Voice Cloning
FULL : voice agent avec voix, contexte, langue, vitesse, style et stabilité configurables. Custom/LITE : apportez votre TTS et votre agent ; c’est la voie à tester pour une voix française et une direction émotionnelle maîtrisées dans Où est Ava ? Le glossaire sert la prononciation sur les surfaces prévues ; il n’est pas une direction de jeu. Le clonage professionnel reste une option d’accès et de consentement à vérifier.
Persona Fine-Tuning
Modular : avatar_id (visuel) + voice_id (audio) + context_id (personnalité) découplés. Un même persona peut garder sa voix/agent sur des visages distincts. La personnalisation visuelle ne crée pas un contrôle de geste live par réplique ; v3 ajoute style_id aux inserts asynchrones.
Avatar Training
Best Practices
- 01.15s listening segment (smiles, nods)
- 02.90s natural speech on any topic
- 03.15s stillness (eyes on camera, natural breathing)
- 04.Uniform lighting, no shadows — background is baked into avatar
- 05.Silent environment, stable position
- 06.Closed mouth pauses between sentences
- 07.For v3 Digital Twin: enable matting if remove_background needed
API Analysis
Protocols
SDKs
Concurrent Sessions
Free : 1 session · Starter, Essential et Business : concurrence illimitée annoncée · Enterprise : capacité et SLA dédiés
Rate Limits
La concurrence n’est plus facturée ni plafonnée sur les plans payants ; crédits, durée maximale de session et éventuels overages restent applicables.
Key Features
- v3 unified API (api.heygen.com) — Video Agent flagship, Avatar V engine, Cinematic Avatar, HyperFrames
- LiveAvatar Custom/LITE : apportez STT, LLM, TTS et orchestration — 1 crédit/min, adapté à la pile GamiWays
- LiveAvatar FULL : ASR, LLM, TTS, voice agent et streaming gérés — 2 crédits/min
- Concurrence illimitée annoncée dès Starter : payer les minutes, pas des slots de session
- Connectors: ElevenLabs Agent, OpenAI Realtime, Gemini Live
- Plugins: LiveKit, Pipecat, Agora, VisionAgents
- Plan gratuit : 10 crédits, 1 session et 2 minutes maximum par session
- Webhooks on v3 (callback_url) and LiveAvatar session events
- MCP + CLI for agent-native workflows
API Constraints
- Two incompatible API key systems (LiveAvatar vs HeyGen Studio/v3)
- Interactive Avatar API sunset March 31, 2026
- Real-time: no scene/staging API — only avatar_id + quality + encoding
- Custom/LITE : LiveKit par défaut ou infrastructure WebRTC LiveKit/Agora propre ; H264 recommandé
- Cinematic Avatar: no voice/script, 4–15s clips, $7 flat
- Avatar V rejects motion_prompt and expressiveness (Avatar IV only)
- v1/v2 legacy endpoints deprecated October 31, 2026
Pricing Model
| Plan | Price | Included minutes | Overage |
|---|---|---|---|
| LiveAvatar Free | $0/mo | 10 crédits · 1 session · 2 min/session | Watermark |
| LiveAvatar Starter | $19/mo | 200 crédits · 5 min/session · concurrence illimitée | $0.10/crédit |
| LiveAvatar Essential | $99/mo | 1 100 crédits · 20 min/session · concurrence illimitée | $0.095/crédit |
| LiveAvatar Business | $475/mo | 6 000 crédits · 60 min/session · concurrence illimitée | $0.09/crédit |
| v3 async | Pay-per-use | N/A (per video/operation) | Cinematic: $7/video flat |
| Enterprise | Custom | Custom | Negotiated |
Hidden costs / watch out
- FULL consomme 2 crédits/min et Custom/LITE 1 crédit/min : comparer les architectures sur une même unité
- Custom/LITE ajoute les coûts de STT, LLM, TTS, RAG, orchestration et WebRTC propres à Où est Ava ?
- L’avatar custom 720p est inclus dans Essential, le custom 1080p dans Business ; la 1080p augmente la latence
- 4K async video billed at premium rate on v3
Sovereignty & Hosting
Sovereignty Score
Hosting
AWS US-East-1
GDPR
YesOn-premise
NoSovereignty detail
AWS US-East only. No EU hosting. GDPR DPA available. Separate API keys for LiveAvatar vs HeyGen Studio/v3.
Constraints & Limits
- Real-time scene/staging NOT programmable — biggest gap vs GamiWays behavioral vision
- Two separate APIs and key systems to manage
- Avatar custom créé depuis une image ou deux minutes de captation ; consentement et qualité de la source à prévoir
- AWS US hosting only — no EU option
- Strict content moderation
- Async v3 (3–8s+) incompatible with conversational latency targets alone
GamiWays Relevance
Score
8/10
À tester en priorité pour Où est Ava ? : LiveAvatar Custom/LITE comme moteur vidéo au-dessus du STT, LLM/RAG et TTS français de GamiWays. Cette séparation permet de faire varier le « comment le dire » (Inworld, Gradium, etc.) sans confondre voix et rendu vidéo. La concurrence illimitée annoncée dès Starter retire un frein de capacité, mais TTFF, première parole, lip-sync, reprises d’interruption, qualité française et coût doivent être mesurés avec la pile complète. Garder les API v3 asynchrones pour les inserts où la mise en scène compte : la scénographie temps réel reste liée au matériau d’entraînement.