The current prototype
What works, measured latency and technical constraints of the current prototype.
What already works
Actual state of the Où est Ava ? prototype — not a product promise.
Web interface, post-film journey, voice capture, multi-provider STT, multi-model LLM, multi-provider TTS and Max or Emma depending on readiness. The prototype combines character-scoped narrative RAG, private memory, Game Master, anti-hallucination validator, per-turn vocal performance intents, admin dashboard, logs, latency metrics and questionnaire. The target remains fictional video calling with a streaming video avatar; wider releases remain conditional on quality, journey and reliability testing.
What remains
- Artistic quality over time
- Complete Ava and Léo
- Fluidity across four viewpoints
- Voice expressivity and stability
- Video embodiment of all four
- Relevance and dosage of cinematics
- Operational stability
- Optimal duration, to be set by tests
- Actual effect on audiences
Game Master
Invisible director.
Artistic target: an invisible director.
Implemented: a relational Game Master after the turn.
Still to calibrate: prompts, cinematics and the arc of a full session.
Latency — narrative challenge
Latency conditions character credibility. Documented ranges and prototype services — PostHog measurements being stabilised below.
Latency is not a simple technical problem: it directly affects suspension of disbelief. Each speech turn crosses a chain of services (STT, query rewrite, RAG, Game Master, LLM, validator, TTS) before Max speaks. The ranges below are documented orders of magnitude — Où est Ava ? design budgets (Notion) and market references — not yet precise measurements from the prototype in production.
Latency budget — orders of magnitude
Cumulative design budgets (excl. STT, excl. avatar): ~2.8 s. Back-office experience target: ~2 s. Minimal voice cascade market reference (Deepgram + GPT-4o + fast TTS): ~500 ms best case.
Sources: Notion Où est Ava ? pipeline (prototypeDiagrams), portal benchmarks (sttData, pipelineData, WhatWeHaveDiagram).
Browse each link in the speech turn. Listed services match the current prototype.
Streaming voice transcription via WebSocket + VAD — first measurable link in an audio turn.
Services in use
- →Gamilab API (proxy-stt, orchestration)
- →Deepgram Nova (primary)
- →Gamilab / Whisper / AssemblyAI (admin alternatives)
Range / budget
Market: ~75–200 ms (Deepgram Nova, streaming P90 / typical)
Narrative stake
Slow transcription delays the entire pipeline; the user already hears silence before the character even 'thinks'.
Tap an item to show its detail below.
Browse each link in the speech turn. Listed services match the current prototype.
Streaming voice transcription via WebSocket + VAD — first measurable link in an audio turn.
Services in use
- →Gamilab API (proxy-stt, orchestration)
- →Deepgram Nova (primary)
- →Gamilab / Whisper / AssemblyAI (admin alternatives)
Range / budget
Market: ~75–200 ms (Deepgram Nova, streaming P90 / typical)
Narrative stake
Slow transcription delays the entire pipeline; the user already hears silence before the character even 'thinks'.
Analytics & sessions
PostHog metrics — latency, journey and session quality in real conditions.
PostHog technical monitoring
Monitoring page to observe performance, objectify improvements and correlate latency with experience quality.