Simli (Trinity-1)
Ultra-low latency real-time avatar from a single image
Latence TTFR
~300ms
temps réel
Coût / minute
à partir de $0.010/min
selon l'abonnement souscrit
Qualité visuelle
7/10
score estimé
Protocoles
WebRTC, REST, WebSocket, LiveKit, Pipecat
Direction émotionnelle & mise en scène
Des états faciaux nommés peuvent être appliqués directement au streaming ; c’est un contrôle explicite mais volontairement discret.
Pilotage API expliciteParamètres et mécanisme
`emotionId` / `emotion_id` pour les presets Natural, Happy, Angry, Doubtful, plus `handleSilence`.
Limite à garder visible
Aucun geste, regard ou caméra par tour n’est documenté ; synchroniser le choix d’état avec une voix expressive.
Personnalisation de l'avatar
RAG / Base de connaissances
Via third-party LLM integration (OpenAI, Anthropic via Pipecat/LiveKit). Custom knowledge bases fed to the LLM layer. Simli handles Speech-to-Video only.
Comportement & personnalité
Behavior defined by connected LLM prompt. Simli is a Speech-to-Video renderer — personality lives in the LLM layer.
Langage corporel & gestes
Head movements and facial micro-expressions auto-generated. No complex hand/body gesture API.
Expressions faciales
Realistic facial expressions and smooth animation via Trinity-1. Gaussian model for photorealistic face cloning.
Voix & clonage vocal
ElevenLabs integration for voice customisation (tone, accent, speed). Simli handles audio-to-video sync.
Fine-tuning du persona
Persona lives in the LLM layer (external). Simli only handles visual rendering from audio input.
Entraînement de l'avatar
Bonnes pratiques
- 01.Front-facing photo, well-lit, neutral expression
- 02.Closed mouth
- 03.No obstructions (glasses, hair over face)
- 04.Gaussian model: stricter quality requirements for photorealism
Analyse API
Protocoles
SDKs
Sessions simultanées
1 (Free) → 2 (Hobby) → 10 (Pro) → 50 (Scale)
Limites de débit
Avatar slots: 1 (Free) → 1 (Hobby) → 5 (Pro) → 30 (Scale)
Fonctionnalités clés
- POST /compose/token — session token
- GET /compose/ice — ICE servers for WebRTC
- Native LiveKit and Pipecat integration
- Speech-to-Video pipeline: audio in → video out
- <300ms end-to-end latency
Contraintes API
- Manual WebRTC negotiation required for custom implementations
- No built-in LLM or TTS (bring your own)
- Avatar slots limited by plan
- No webhook support
Modèle tarifaire
| Plan | Prix | Minutes incluses | Dépassement |
|---|---|---|---|
| Free | $0/mo | 50 min/mo | N/A |
| Hobby | $10/mo | 1000 min/mo | $0.01/min |
| Pro | $49/mo | 5500 min/mo | $0.0095/min |
| Scale | $249/mo | 27500 min/mo | $0.009/min |
| Enterprise | Custom | Custom | Custom |
Coûts cachés / à surveiller
- ElevenLabs TTS billed separately
- LLM API costs (OpenAI/Anthropic) billed separately
Souveraineté & Hébergement
Score de souveraineté
Hébergement
Cloud (Norwegian company, EU jurisdiction)
GDPR
YesOn-premise
NoDétail souveraineté
Norwegian company (Simli AS). EU jurisdiction. No explicit EU datacenter confirmed. No on-premise.
Contraintes & Limites
- No body gesture API (head movements only)
- No built-in LLM or TTS — must integrate separately
- Avatar slots limited by plan tier
- No webhook support
- No on-premise option
Pertinence pour GamiWays
Score
9/10
Best price/performance ratio for real-time video rendering. Ideal as a Speech-to-Video module in GamiWays's modular pipeline. Ultra-low cost ($0.009/min) and <300ms latency. Limitation: no built-in AI stack — must integrate ASR/LLM/TTS separately.