StepAudio 2.5 TTS
Contextual TTS — ELO 1187, dual-level context control, zero-shot voice cloning in 3 sec
Comparative Scores
Architecture
Exceptional quality for storytelling and emotional content (ELO #3). The dual-level context control is a unique differentiator for GamiWays character performance. However, Chinese jurisdiction and no GDPR compliance are blocking factors for Swiss/EU clients. Recommended for quality benchmarking and non-sensitive use cases only.
Analysis
StepAudio 2.5 TTS (launched April 16, 2026) ranks #3 globally on Artificial Analysis Speech Arena (ELO 1187). It is the first model to integrate contextual understanding into the entire speech generation pipeline. Dual-level context control: Global Context sets the overall tone for a passage; Inline Context uses parentheses for per-sentence fine-grained control of emotion, pauses, and breathing. Zero-shot voice cloning from 3 seconds of reference audio. 100+ languages. Price: $0.85/10K chars ($0.064/min). step-tts-2 (standard): $0.40/10K chars ($0.030/min). Chinese company (StepFun, Shanghai) — no GDPR/sovereignty.
Strengths
- ELO 1187 — rank #3 globally (Artificial Analysis)
- Dual-level context control (Global + Inline)
- Zero-shot voice cloning from 3 sec reference
- 100+ languages with crosslingual identity
- step-tts-2 at $0.030/min — competitive pricing
Weaknesses
- Chinese company — no GDPR, no EU data residency
- No on-premise option
- $85/1M chars (flagship) — expensive vs Inworld ($42/1M)
- No lip-sync timestamps
- Max 1000 chars per request
Voice Capabilities
Zero-shot voice cloning from just 3 seconds of reference audio. Full Global/Inline context control inherited by cloned voice. $1.50/voice.
A global instruction and per-sentence parentheses separate performance intent from spoken text.
How: `instruction`, `voice`, `input` and inline context such as `(restrained, lowered voice)`.
Validate: Verify language support, compliance and consistency before sensitive use in Où est Ava ?
WebSocket streaming via /v1/realtime/audio. Non-streaming POST /v1/audio/speech also available.
No native lip-sync timestamps. External aligner required.
Pricing
$0.85/10K chars ($85/1M chars). Voice cloning: $1.50/voice. step-tts-2 (standard): $0.40/10K chars ($40/1M chars).
Sovereignty & Compliance
Cloud-only. StepFun is a Shanghai-based company (Chinese jurisdiction).
Data residency: China (StepFun servers, Shanghai). No EU data residency option.
This sheet's verification
Artificial Analysis Speech Arena, May 2026
Update note: AA Speech Arena sync 2026-08-18: ELO 1205 (rank N/A — Free tier). Name: StepAudio 2.5 TTS.
This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.
ELO benchmarks / indices: Artificial Analysis · Last API sync : 18 August 2026