GamiWays
Cloud API#10 Artificial AnalysisCommercial

MiniMax Speech 2.8

Sound tags natifs — $0.10/1M chars, 17 langues, clonage vocal instant

200 ms
TTFA (best case)
400 ms
TTFA (typical)
$0.1/1M
Price per million chars
1175
ELO Score

Comparative Scores

Voice quality8/10
Latency6/10
Voice cloning8/10
Expressiveness9/10
Performance direction10/10
Sovereignty1/10
Price accessibility10/10
Multilingual7/10

Architecture

ArchitectureProprietary TTS (MiniMax, Shanghai)
ParametersUndisclosed
Languages17
Self-hostable No
Streaming Yes
GamiWays
Prototypage économique — non recommandé pour production EU

Pertinent pour le prototypage Storygami/Edugami à faible coût. Les sound tags natifs sont un différenciateur unique pour les personnages expressifs. Bloqué pour la production EU (juridiction chinoise, pas de RGPD). Recommandé pour : benchmarking expressivité, prototypage rapide, cas d'usage non-sensibles.

Analysis

MiniMax Speech 2.8 is the most affordable TTS API on the market at $0.10/1M chars. Native sound tags ([laugh], [sigh], [cry], [surprised], etc.) enable para-language control without SSML complexity. 17 languages, instant voice cloning at same price. Chinese company (Shanghai) — no GDPR compliance. Best for cost-sensitive prototyping and non-EU deployments.

Strengths

  • $0.10/1M chars — prix le plus bas du marché
  • Sound tags natifs : [laugh], [sigh], [cry], [surprised], etc.
  • Clonage vocal instant au même prix
  • 17 langues
  • Très compétitif pour le prototypage

Weaknesses

  • Entreprise chinoise — pas de RGPD, pas de résidence EU
  • Pas d'option on-premise
  • Pas de timestamps lip-sync
  • Pas de score ELO (non classé Artificial Analysis)

Voice Capabilities

Voice Cloning Yes

Instant voice cloning. 17+ languages. Voice cloning at same price as standard TTS.

Emotion & performance directionExplicit API control

Prosody and non-verbals can be combined to steer a line toward a precise emotion and rhythm.

How: `voice_setting` (`speed`, `pitch`, `vol`), `sound_effects`, pauses and `(laughs)` / `(sighs)` tags.

Validate: Verify French quality, rights, jurisdiction and tag interpretation before any decision.

Streaming Yes

Streaming WebSocket. Sub-500ms TTFA.

Lip-sync Data No

No native lip-sync timestamps.

Pricing

Price / 1M chars
from $0.1
depending on plan
Price / minute
from $0.0001
depending on plan
Free tier
Free tier available

$0.10/1M chars (Speech 2.8). Speech 2.0 HD: $0.60/1M chars. Streaming: $0.10/1M chars. Clonage vocal: $0.10/1M chars. Très compétitif.

Sovereignty & Compliance

On-premise No

Cloud only. Chinese jurisdiction (MiniMax, Shanghai).

GDPR No

Data residency: China (MiniMax servers, Shanghai). No EU data residency.

This sheet's verification

Verified 18 August 2026

MiniMax platform docs + coval.ai TTS comparison 2026

Update note: AA Speech Arena sync 2026-08-18: ELO 1175 (rank N/A — Free tier). Name: Speech 2.8 HD.

This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.

ELO benchmarks / indices: Artificial Analysis · Last API sync : 18 August 2026