GamiWays
Cloud API#5 Artificial AnalysisProprietary (StepFun / 阶跃星辰)

StepAudio 2.5 TTS

Contextual TTS — ELO 1187, dual-level context control, zero-shot voice cloning in 3 sec

200 ms
TTFA (best case)
400 ms
TTFA (typical)
$85/1M
Price per million chars
1205
ELO Score

Comparative Scores

Voice quality9/10
Latency6/10
Voice cloning9/10
Expressiveness10/10
Performance direction7/10
Sovereignty1/10
Price accessibility5/10
Multilingual9/10

Architecture

ArchitectureEnd-to-end contextual TTS (NTP-based, no speaker/emotion embeddings)
ParametersUndisclosed
Languages100
Self-hostable No
Streaming Yes
GamiWays
Benchmarking qualité — non recommandé pour production EU

Exceptional quality for storytelling and emotional content (ELO #3). The dual-level context control is a unique differentiator for GamiWays character performance. However, Chinese jurisdiction and no GDPR compliance are blocking factors for Swiss/EU clients. Recommended for quality benchmarking and non-sensitive use cases only.

Analysis

StepAudio 2.5 TTS (launched April 16, 2026) ranks #3 globally on Artificial Analysis Speech Arena (ELO 1187). It is the first model to integrate contextual understanding into the entire speech generation pipeline. Dual-level context control: Global Context sets the overall tone for a passage; Inline Context uses parentheses for per-sentence fine-grained control of emotion, pauses, and breathing. Zero-shot voice cloning from 3 seconds of reference audio. 100+ languages. Price: $0.85/10K chars ($0.064/min). step-tts-2 (standard): $0.40/10K chars ($0.030/min). Chinese company (StepFun, Shanghai) — no GDPR/sovereignty.

Strengths

  • ELO 1187 — rank #3 globally (Artificial Analysis)
  • Dual-level context control (Global + Inline)
  • Zero-shot voice cloning from 3 sec reference
  • 100+ languages with crosslingual identity
  • step-tts-2 at $0.030/min — competitive pricing

Weaknesses

  • Chinese company — no GDPR, no EU data residency
  • No on-premise option
  • $85/1M chars (flagship) — expensive vs Inworld ($42/1M)
  • No lip-sync timestamps
  • Max 1000 chars per request

Voice Capabilities

Voice Cloning Yes

Zero-shot voice cloning from just 3 seconds of reference audio. Full Global/Inline context control inherited by cloned voice. $1.50/voice.

Emotion & performance directionInstruction-based direction

A global instruction and per-sentence parentheses separate performance intent from spoken text.

How: `instruction`, `voice`, `input` and inline context such as `(restrained, lowered voice)`.

Validate: Verify language support, compliance and consistency before sensitive use in Où est Ava ?

Streaming Yes

WebSocket streaming via /v1/realtime/audio. Non-streaming POST /v1/audio/speech also available.

Lip-sync Data No

No native lip-sync timestamps. External aligner required.

Pricing

Price / 1M chars
from $85
depending on plan
Price / minute
from $0.0640
depending on plan
Free tier
No public free tier. API key required.

$0.85/10K chars ($85/1M chars). Voice cloning: $1.50/voice. step-tts-2 (standard): $0.40/10K chars ($40/1M chars).

Sovereignty & Compliance

On-premise No

Cloud-only. StepFun is a Shanghai-based company (Chinese jurisdiction).

GDPR No

Data residency: China (StepFun servers, Shanghai). No EU data residency option.

This sheet's verification

Verified 18 August 2026

Artificial Analysis Speech Arena, May 2026

Update note: AA Speech Arena sync 2026-08-18: ELO 1205 (rank N/A — Free tier). Name: StepAudio 2.5 TTS.

This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.

ELO benchmarks / indices: Artificial Analysis · Last API sync : 18 August 2026