GamiWays

Gradium.ai

New

TTS, STT, Voice Cloning & S2S Translation — ultra-low latency, academic founders

TTSSTTVoice CloningS2S TranslationOn-devicePhononLiveKitPipecatFrançais

01Overview

Gradium develops audio language models designed to deliver natural, expressive, ultra-low latency voice interactions at scale. Its founders have collectively invented and published the methods and algorithms behind most voice models existing today — from neural audio codecs to audio language models.

The company translates more than a decade of open research into production-ready systems. Gradium offers TTS, STT, Voice Cloning, and a unique real-time Speech-to-Speech Translation capability. In July 2026, Gradium raised $100M (NVIDIA investor) and opened a San Francisco office, while launching Phonon Multilingual — a 100M-parameter on-device model covering 5 languages (FR/EN/DE/ES/PT), running offline on CPU.

02Products & Capabilities

TTS — Text-to-Speech
  • ·Streaming and word-level timestamps for avatar lip-sync
  • ·216ms P50 TTFA and 30ms IQR reported on Coval
  • ·81% on structured entities in five languages including French (vendor benchmark)
  • ·Voice cloning according to plan, consent and zone to verify
  • ·Mid-sentence multilingual code-switching
  • ·Phonon on-device is distinct from the cloud API
STT — Speech-to-Text
  • ·Best-in-class accuracy (claimed)
  • ·Semantic VAD — intelligent turn detection
  • ·Noise robustness
  • ·Controllable latency per use case
  • ·Speaker diarization
  • ·Word-level timestamps
Gradium Translate — S2S
  • ·Real-time Speech-to-Speech Translation
  • ·Speech-to-Text Translation
  • ·5 languages: EN, FR, ES, DE, PT
  • ·Best BLEU/MetricX vs Gemini and GPT-Realtime (pareto)
  • ·Bidirectional WebSocket API
  • ·Automatic source language detection

Supported languages : English, French, Spanish, German, Brazilian Portuguese — multilingual mid-sentence code-switching without additional latency.

02b2026 Product markers

Jul 8, 2026

$100M Funding + NVIDIA

Seed round extended to $100M total. NVIDIA among investors. New San Francisco Bay Area office opened.

Jul 15, 2026

Phonon Multilingual

On-device model, 100M params, 5 languages (FR/EN/DE/ES/PT). WER: FR 2.18%, DE 0.50%, ES 0.53%, PT 1.31%. Beats NVIDIA Magpie (357M) at 3.6× fewer parameters. Runs on CPU, ~200 MB, offline.

Aug 31, 2026

New TTS model

Gradium reports 216ms P50 TTFA and 30ms IQR on Coval, plus 81% on 500 structured-entity phrases across five languages including French. These vendor metrics do not evaluate emotion, network or video avatar.

Jul 2026

Semantic VAD + GradBot

Semantic turn detection (predictions every 80 ms). GradBot: open-source voice agent framework in 50 lines of code (Rust core).

03Key Differentiators

Academic founders

Inventors of neural audio codecs and audio language models. Over a decade of open research → production.

Semantic VAD

Intelligent turn detection based on meaning, not just silence. Critical for natural conversational agents.

Unique S2S Translation

Real-time Speech-to-Speech Translation with best BLEU/MetricX vs Gemini and GPT-Realtime at low latency.

Native LiveKit & Pipecat integration

Plug-and-play in the most widely used voice agent frameworks. No custom integration needed.

Residency and on-device path

EU or US residency must be enabled then verified on a paid plan; on-device Phonon is a separate path. A French company does not guarantee Paris hosting or automatic sovereignty.

Usage-based pricing

Flexible credit system. Free tier without credit card. XS plan at $13/month sufficient for prototyping.

04Pricing

Credit rates per service

TTS

1 crédit / caractère

1h TTS ≈ 45k crédits

STT

3 crédits / seconde

1h STT = 10,800 crédits

STT Translation

4 crédits / seconde

1h = 14,400 crédits

S2S Translation

30 crédits / seconde

1h = 108,000 crédits

Monthly plans

PlanPrice/moCreditsTTS hoursSTT hoursConc. TTSConc. STT
FreeNo CC$045k~1h4h23
XS$13225k~5h21h520
S$43900k~20h83h520
M$3409M~200h833h1040
L$1,61545M~1,000h4,167h1560
EnterpriseCustom∞∞∞CustomCustom

Additional credits: $3.8–$6.9 / 100k credits depending on plan. TTS concurrency: 2/5/5/10/15/Custom. STT concurrency: 3/20/20/40/60/Custom. Source: gradium.ai/pricing (July 2026).

05Integrations & Tech Stack

Voice Agent Frameworks

LiveKit AgentsNative integration — plug-and-play
PipecatNative integration — declarative pipeline

SDKs & API

Python SDKOfficial client
Rust SDKOfficial client
WebSocket APIBidirectional, streaming
REST APIBatch and async

Deployment

Cloud APIAll plans
EU/US residencyPaid plan · enable + verify
Phonon on-deviceSeparate from the cloud API

Sovereignty & GDPR

· EU or US residency: paid plan, activation and response signal to verify

· On-device Phonon: separate path; do not infer its guarantees from the cloud API

· ZDR, cloned voices and any on-premise terms: separate contractual scope

· A French company ≠ Paris hosting or automatic Swiss sovereignty

06Relevance for Storygami & Edugami

1

STT for Storygami & Edugami

Semantic VAD → natural turn detection for AI characters. FR/EN/DE/ES/PT languages covered. Noise robustness in classroom or cinema settings.

2

Precise TTS + lip-sync

Word-level timestamps and structured-entity pronunciation serve avatar lip-sync and French terminology. Où est Ava ? protagonists’ and Edugami mentor’s emotional prosody still needs testing; no explicit control is documented.

3

Multilingual S2S Translation

Unique use case: French-speaking student talks, avatar responds in English (or vice versa). Relevant for Edugami in multilingual school contexts.

4

Real-Time Infrastructure integration

Natively compatible with LiveKit and Pipecat. Aligned with the portal's Real-Time Infrastructure roadmap. Possible evaluation in GamiWays Core Phase B.

5

Residency & Edugami

Enable the EU endpoint on a paid plan then verify its signal. EU residency is neither Swiss residency nor an automatic guarantee for cloned voice assets; evaluate it against the intended school framework.

07Open Questions

  • ?What is the actual TTFA latency of Gradium TTS from Switzerland? (benchmark to perform)
  • ?Is Gradium STT WER independently validated (Pipecat benchmark, Artificial Analysis)?
  • ?Is EU residency enabled on the selected plan, and does the response signal actually confirm the zone?
  • ?Are API pauses and prosody credible enough for Où est Ava ? despite the absence of explicit emotional direction?
  • ?Is S2S Translation production-stable (latency, accuracy) for school use in FR→EN?

Related pages

Sources

Data verified: July 2026.