GamiWays
Open SourceModified MIT (open-source weights on Hugging Face)

Miso-TTS v1 (8B)

110ms latency — most emotive open-source TTS with RVQ architecture

110 ms
TTFA (best case)
110 ms
TTFA (typical)
Free
Price per million chars
—
ELO Score

Comparative Scores

Voice quality9/10
Latency9/10
Voice cloning9/10
Expressiveness10/10
Performance direction5/10
Sovereignty10/10
Price accessibility10/10
Multilingual1/10

Architecture

ArchitectureHierarchical RVQ Transformer (7.7B backbone + 300M depth decoder)
Parameters8B (7.7B backbone + 300M depth transformer)
Languages1
Self-hostable Yes
Streaming No
GamiWays
Souveraineté + Expressivité — Prototype self-hosted

Highly relevant for GamiWays sovereignty requirements: open-source weights, on-premises deployment, 110ms latency. The audio conditioning on user tone is a unique differentiator for emotionally responsive AI characters. Blocked by: half-duplex only (no interruption handling), English-focused, no cloud API yet. Recommended for: self-hosted prototype evaluation, sovereignty-first deployments, emotional expressiveness benchmarking.

Analysis

Miso-TTS v1 is an 8B-parameter open-source TTS model released June 3, 2026. It achieves 110ms latency — among the lowest publicly claimed on the market — using a hierarchical RVQ transformer architecture (7.7B backbone + 300M depth decoder). The model conditions on both text and the user's audio context, enabling emotionally aware speech generation. Open-source weights under modified MIT license. Cloud API coming soon. Half-duplex only; full-duplex is future work.

Strengths

  • 110ms latency — among lowest claimed on market
  • Audio conditioning on user's tone (emotional awareness)
  • Open-source weights (modified MIT) — full sovereignty
  • On-premises enterprise support available
  • RVQ architecture: 2048^32 addressable audio tokens
  • One-shot voice cloning (10s clip)

Weaknesses

  • Cloud API not yet available (coming soon)
  • Half-duplex only — no full-duplex or turn-taking
  • English-focused (multilingual not documented)
  • No ELO score yet (not on Artificial Analysis leaderboard)
  • No pricing for cloud API
  • No lip-sync timestamps

Voice Capabilities

Voice Cloning Yes

One-shot voice cloning from a 10-second audio clip. Exact voice replica from first to last second of a call.

Emotion & performance directionAudio-context steering

Audio context can convey performance colour; no isolated emotional intention API is exposed.

How: `context`, `speaker`, `temperature`, `topk` in open-weight inference.

Validate: Treat as a self-hosted R&D path, with per-state audio references and an evaluation protocol.

Streaming No

Streaming not documented. Half-duplex only — full-duplex and turn-taking are future work.

Lip-sync Data No

No lip-sync timestamps documented.

Pricing

Price / 1M chars
Free
Price / minute
Free
Free tier
Self-hosted (open-source weights, modified MIT license)

Self-hosted: free (compute costs only). Cloud API: coming soon — pricing not yet announced.

Sovereignty & Compliance

On-premise Yes

Open-source weights (modified MIT). On-premises hosting + support contracts available for enterprise teams.

GDPR No

Data residency: Self-hosted: full data sovereignty. Cloud API: not yet available.

RVQ Architecture — Detail

Miso-TTS v1 uses a two-stage hierarchical RVQ Transformer: a 7.7B backbone generates coarse audio tokens, then a 300M depth transformer refines fine details. The audio space is represented by 2048^32 addressable tokens (RVQ depth=32, codebook size=2048), yielding a theoretical expressiveness of 10^105 combinations.

Backbone
7.7B params
Generates coarse tokens
Depth Decoder
300M params
Refines fine details
RVQ Codebook
2048^32
Addressable audio tokens

API Status

The cloud API is announced but not yet available (June 2026). Open-source weights are available on Hugging Face under a modified MIT license. On-premises deployment is available for enterprise teams.

This sheet's verification

Verified 5 June 2026

Miso Labs blog post, June 3, 2026 — internal latency comparison (ElevenLabs 700ms, Sesame 300ms, Human 160ms, Miso 110ms)

Update note: Initial entry from misolabs.ai and blog post (June 3, 2026). No AA ELO yet — API not yet available.

This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.