Technology
Software stack, Game Master, narrative RAG and technical challenges.
The stack separates interface, voice, characters, orchestration, sources and observation.
The AVA-specific prototype and the generic Core are developed in parallel.
Prototype execution architecture
From speech to delivery: the path that must answer now, then what prepares and measures the next turn.
A web experience, an editing and orchestration application, AI services and content sources connected through clearly separated flows.
Web experience
AI services · listening
Editing & orchestration
Narrative memory
Orchestration
AI services · response
Beside the audible path
Triggered toward the experience, separately from the avatar.
Latency and traces, without delaying the response.
Diagram focus
Conversation engine
This engine runs the live turn. It requests useful context, calls the response model and sends only final text to delivery.
Software stack
17 connected components — logos, roles, voice flow and links between tools.
Three connected layers: where we write (Notion), how the experience is orchestrated (API, RAG, GM), which external services listen and speak (Deepgram, OpenRouter, TTS…). Click a component for detail and links.
Creation / editing tools
Application / orchestration
Application / orchestration
Gamilab API
Technical conductor — chains STT, RAG, GM, LLM, TTS per turn
The API is not a simple proxy: it is the orchestrator deciding call order (query rewrite, scoped RAG retrieval, streaming Max generation, anti-hallucination validation, sentence-level TTS). On the realtime path, pre-turn GM was removed to save seconds; post-turn GM runs async and logs to the session. This is the orchestration core of Gamilab's infrastructure.
What this component does
- Edge functions: proxy-llm, proxy-stt, proxy-tts, query-rag, rewrite-query…
- Aggregates per-segment latencies for admin and PostHog.
- Persists sessions, conversation_log, voice_turn_events.
In the turn flow
Incoming turn: STT → rewrite-query → RAG → Max LLM → validator → TTS → post-turn GM (async)
Links to other components
The layer in one sentence
Layer visible to users and the team: post-film experience in the browser, API chaining AI services, control back-office, and narrative memory feeding Max.
Infrastructure / services
Tap an item to show its detail below.
Three connected layers: where we write (Notion), how the experience is orchestrated (API, RAG, GM), which external services listen and speak (Deepgram, OpenRouter, TTS…). Click a component for detail and links.
Creation / editing tools
Application / orchestration
Infrastructure / services
Application / orchestration
Gamilab API
Technical conductor — chains STT, RAG, GM, LLM, TTS per turn
The API is not a simple proxy: it is the orchestrator deciding call order (query rewrite, scoped RAG retrieval, streaming Max generation, anti-hallucination validation, sentence-level TTS). On the realtime path, pre-turn GM was removed to save seconds; post-turn GM runs async and logs to the session. This is the orchestration core of Gamilab's infrastructure.
What this component does
- Edge functions: proxy-llm, proxy-stt, proxy-tts, query-rag, rewrite-query…
- Aggregates per-segment latencies for admin and PostHog.
- Persists sessions, conversation_log, voice_turn_events.
In the turn flow
Incoming turn: STT → rewrite-query → RAG → Max LLM → validator → TTS → post-turn GM (async)
Links to other components
The layer in one sentence
Layer visible to users and the team: post-film experience in the browser, API chaining AI services, control back-office, and narrative memory feeding Max.
Game Master & narrative RAG
Three angles — why a GM, its concrete role, and narrative RAG risks.
An AI character cannot simply be a "well-prompted" LLM. It must respect what it knows, what it ignores, what it hides, its emotional state and the rhythm of the experience. The Game Master acts as an invisible director.
The difficulty is not just making Max speak. The difficulty is preventing him from knowing too much, answering too well, or leaving his world.
Main technical challenges
Seven structuring tensions — browse the list for each challenge in detail.
Seven tensions structuring development — browse each challenge to see action levers.
Latency
Critical challenge — target < 1.5s end-to-end
Action levers
- Fast STT (< 300ms)
- Reactive LLM (streaming)
- Smooth TTS (TTFA < 400ms)
- Visual feedback while waiting
Tap an item to show its detail below.
Seven tensions structuring development — browse each challenge to see action levers.
Latency
Critical challenge — target < 1.5s end-to-end
Action levers
- Fast STT (< 300ms)
- Reactive LLM (streaming)
- Smooth TTS (TTFA < 400ms)
- Visual feedback while waiting