Miso-TTS v1 (8B)
110ms latency — most emotive open-source TTS with RVQ architecture
Comparative Scores
Architecture
Highly relevant for GamiWays sovereignty requirements: open-source weights, on-premises deployment, 110ms latency. The audio conditioning on user tone is a unique differentiator for emotionally responsive AI characters. Blocked by: half-duplex only (no interruption handling), English-focused, no cloud API yet. Recommended for: self-hosted prototype evaluation, sovereignty-first deployments, emotional expressiveness benchmarking.
Analysis
Miso-TTS v1 is an 8B-parameter open-source TTS model released June 3, 2026. It achieves 110ms latency — among the lowest publicly claimed on the market — using a hierarchical RVQ transformer architecture (7.7B backbone + 300M depth decoder). The model conditions on both text and the user's audio context, enabling emotionally aware speech generation. Open-source weights under modified MIT license. Cloud API coming soon. Half-duplex only; full-duplex is future work.
Strengths
- 110ms latency — among lowest claimed on market
- Audio conditioning on user's tone (emotional awareness)
- Open-source weights (modified MIT) — full sovereignty
- On-premises enterprise support available
- RVQ architecture: 2048^32 addressable audio tokens
- One-shot voice cloning (10s clip)
Weaknesses
- Cloud API not yet available (coming soon)
- Half-duplex only — no full-duplex or turn-taking
- English-focused (multilingual not documented)
- No ELO score yet (not on Artificial Analysis leaderboard)
- No pricing for cloud API
- No lip-sync timestamps
Voice Capabilities
One-shot voice cloning from a 10-second audio clip. Exact voice replica from first to last second of a call.
Audio context can convey performance colour; no isolated emotional intention API is exposed.
How: `context`, `speaker`, `temperature`, `topk` in open-weight inference.
Validate: Treat as a self-hosted R&D path, with per-state audio references and an evaluation protocol.
Streaming not documented. Half-duplex only — full-duplex and turn-taking are future work.
No lip-sync timestamps documented.
Pricing
Self-hosted: free (compute costs only). Cloud API: coming soon — pricing not yet announced.
Sovereignty & Compliance
Open-source weights (modified MIT). On-premises hosting + support contracts available for enterprise teams.
Data residency: Self-hosted: full data sovereignty. Cloud API: not yet available.
RVQ Architecture — Detail
Miso-TTS v1 uses a two-stage hierarchical RVQ Transformer: a 7.7B backbone generates coarse audio tokens, then a 300M depth transformer refines fine details. The audio space is represented by 2048^32 addressable tokens (RVQ depth=32, codebook size=2048), yielding a theoretical expressiveness of 10^105 combinations.
API Status
The cloud API is announced but not yet available (June 2026). Open-source weights are available on Hugging Face under a modified MIT license. On-premises deployment is available for enterprise teams.
This sheet's verification
Miso Labs blog post, June 3, 2026 — internal latency comparison (ElevenLabs 700ms, Sesame 300ms, Human 160ms, Miso 110ms)
Update note: Initial entry from misolabs.ai and blog post (June 3, 2026). No AA ELO yet — API not yet available.
This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.