GamiWays
Cloud API#17 Artificial AnalysisCommercial

Fish Audio OpenAudio S1

Pay-as-you-go voice cloning — 70% cheaper than ElevenLabs

200 ms
TTFA (best case)
400 ms
TTFA (typical)
$15/1M
Price per million chars
1124
ELO Score

Comparative Scores

Voice quality6/10
Latency6/10
Voice cloning8/10
Expressiveness7/10
Performance direction7/10
Sovereignty4/10
Price accessibility7/10
Multilingual5/10

Architecture

ArchitectureFlow matching (proprietary)
ParametersN/A (cloud)
Languages13
Self-hostable No
Streaming Yes
GamiWays
Phase 1 MVP — Coût/Souveraineté

Good cost/quality ratio for Phase 1 MVP. S1-mini self-hosting option aligns with sovereignty requirements. Voice cloning included without extra cost is a significant advantage.

Analysis

Fish Audio OpenAudio S1 ranks #6 on Artificial Analysis (ELO 1128, May 2026). Voice cloning included at the same price as basic TTS — no extra fees. 70% cheaper than ElevenLabs. S1-mini open-source weights enable self-hosting for sovereign deployments.

Strengths

  • ELO 1128 — rank #6
  • Voice cloning included at no extra cost
  • 70% cheaper than ElevenLabs
  • S1-mini open-source for self-hosting
  • 13 languages

Weaknesses

  • ~200ms TTFA (slower than Cartesia)
  • No native lip-sync data
  • Limited documentation

Voice Capabilities

Voice Cloning Yes

Voice cloning included at no extra cost. 10-second audio sample. 70% cheaper than ElevenLabs for equivalent quality.

Emotion & performance directionInstruction-based direction

Large repertoire of emotional and non-verbal tags, complemented by speed and volume prosody.

How: `prosody` object (`speed`, `volume`), `[happy]`, `[sad]`, `[whispering]`, `[laughing]`, `[break]` tags and sampling parameters.

Validate: Test tag stability over WebSocket and linguistic relevance before using it in a sensitive scene.

Streaming Yes

Streaming API available. ~200ms TTFA.

Lip-sync Data No

No native lip-sync timestamps.

Pricing

Price / 1M chars
from $15
depending on plan
Price / minute
from $0.0150
depending on plan
Free tier
Limited free tier

$15/1M chars. Pay-as-you-go, no subscription. Voice cloning included at same price.

Sovereignty & Compliance

On-premise No

Cloud API. S1-mini weights available for self-hosting.

GDPR Compliant

Data residency: US/Asia

Strategic & Business Analysis

Fish Audio OpenAudio S1 — Strategic Positioning

Beyond technical specs: where does this tool sit in the ecosystem, what are the risks and strategic implications for GamiWays?

Fish Audio is the open-source disruptor: enterprise-grade voice cloning with natural language emotion control, available both as a cloud API and a self-hostable model — the clearest path from Phase 1 speed to Phase 2 sovereignty.

Cloud + On-premise
Lock-in risk:Low
Sovereignty fit:High
Open-source threat:Low
Pricing:Commoditizing ↓↓

A. Strategic Positioning

Target customer: Developer / SMB / Enterprise — voice cloning, multilingual content

Open-source S2 model with cloud API — enterprise-grade voice cloning and expressiveness at a fraction of ElevenLabs cost.

B. Competitive Moat

  • Open-source S2 model (self-hostable) + cloud API — dual deployment flexibility
  • Natural language emotion tags — superior expressiveness control vs SSML
  • 50% enterprise revenue share — strong institutional adoption signal

Vulnerability: No explicit compliance certifications (SOC2, HIPAA). Fast-moving open-source market could be disrupted by better-funded competitors.

E. Strategic Questions for GamiWays

Sovereignty fit

Open-source S2 model enables full self-hosting in Swiss/EU infrastructure. No compliance certs for cloud API, but self-hosted path is clear.

Build vs. Buy

Build (self-host S2) for Phase 2 sovereignty. Buy (cloud API) for Phase 1 speed. Best of both worlds.

Lock-in risk

Open-source S2 model eliminates vendor lock-in. Cloud API has moderate lock-in but self-hosted alternative always available.

Roadmap alignment

Excellent: cloud API for Phase 1 speed, self-hosted S2 for Phase 2 sovereignty. Natural migration path.

This sheet's verification

Verified 18 August 2026

Artificial Analysis Speech Leaderboard + Fish Audio docs, 2026

Update note: ELO 1124 du 18 août conservé (nom AA : Fish Audio S2 Pro). Le 30 septembre, aucun article Fish n’a été relu pour promouvoir S2.1 Pro : la fiche reste OpenAudio S1. Les labels d’émotion STT ne sont pas ajoutés.

This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.

ELO benchmarks / indices: Artificial Analysis · Last API sync : 18 August 2026