Gradium.ai
NewTTS, STT, Voice Cloning & S2S Translation — ultra-low latency, academic founders
01Overview
Gradium develops audio language models designed to deliver natural, expressive, ultra-low latency voice interactions at scale. Its founders have collectively invented and published the methods and algorithms behind most voice models existing today — from neural audio codecs to audio language models.
The company translates more than a decade of open research into production-ready systems. Gradium offers TTS, STT, Voice Cloning, and a unique real-time Speech-to-Speech Translation capability. In July 2026, Gradium raised $100M (NVIDIA investor) and opened a San Francisco office, while launching Phonon Multilingual — a 100M-parameter on-device model covering 5 languages (FR/EN/DE/ES/PT), running offline on CPU.
02Products & Capabilities
- ·Streaming and word-level timestamps for avatar lip-sync
- ·216ms P50 TTFA and 30ms IQR reported on Coval
- ·81% on structured entities in five languages including French (vendor benchmark)
- ·Voice cloning according to plan, consent and zone to verify
- ·Mid-sentence multilingual code-switching
- ·Phonon on-device is distinct from the cloud API
- ·Best-in-class accuracy (claimed)
- ·Semantic VAD — intelligent turn detection
- ·Noise robustness
- ·Controllable latency per use case
- ·Speaker diarization
- ·Word-level timestamps
- ·Real-time Speech-to-Speech Translation
- ·Speech-to-Text Translation
- ·5 languages: EN, FR, ES, DE, PT
- ·Best BLEU/MetricX vs Gemini and GPT-Realtime (pareto)
- ·Bidirectional WebSocket API
- ·Automatic source language detection
Supported languages : English, French, Spanish, German, Brazilian Portuguese — multilingual mid-sentence code-switching without additional latency.
02b2026 Product markers
$100M Funding + NVIDIA
Seed round extended to $100M total. NVIDIA among investors. New San Francisco Bay Area office opened.
Phonon Multilingual
On-device model, 100M params, 5 languages (FR/EN/DE/ES/PT). WER: FR 2.18%, DE 0.50%, ES 0.53%, PT 1.31%. Beats NVIDIA Magpie (357M) at 3.6× fewer parameters. Runs on CPU, ~200 MB, offline.
New TTS model
Gradium reports 216ms P50 TTFA and 30ms IQR on Coval, plus 81% on 500 structured-entity phrases across five languages including French. These vendor metrics do not evaluate emotion, network or video avatar.
Semantic VAD + GradBot
Semantic turn detection (predictions every 80 ms). GradBot: open-source voice agent framework in 50 lines of code (Rust core).
03Key Differentiators
Academic founders
Inventors of neural audio codecs and audio language models. Over a decade of open research → production.
Semantic VAD
Intelligent turn detection based on meaning, not just silence. Critical for natural conversational agents.
Unique S2S Translation
Real-time Speech-to-Speech Translation with best BLEU/MetricX vs Gemini and GPT-Realtime at low latency.
Native LiveKit & Pipecat integration
Plug-and-play in the most widely used voice agent frameworks. No custom integration needed.
Residency and on-device path
EU or US residency must be enabled then verified on a paid plan; on-device Phonon is a separate path. A French company does not guarantee Paris hosting or automatic sovereignty.
Usage-based pricing
Flexible credit system. Free tier without credit card. XS plan at $13/month sufficient for prototyping.
04Pricing
Credit rates per service
TTS
1 crédit / caractère
1h TTS ≈ 45k crédits
STT
3 crédits / seconde
1h STT = 10,800 crédits
STT Translation
4 crédits / seconde
1h = 14,400 crédits
S2S Translation
30 crédits / seconde
1h = 108,000 crédits
Monthly plans
| Plan | Price/mo | Credits | TTS hours | STT hours | Conc. TTS | Conc. STT |
|---|---|---|---|---|---|---|
| FreeNo CC | $0 | 45k | ~1h | 4h | 2 | 3 |
| XS | $13 | 225k | ~5h | 21h | 5 | 20 |
| S | $43 | 900k | ~20h | 83h | 5 | 20 |
| M | $340 | 9M | ~200h | 833h | 10 | 40 |
| L | $1,615 | 45M | ~1,000h | 4,167h | 15 | 60 |
| Enterprise | Custom | ∞ | ∞ | ∞ | Custom | Custom |
Additional credits: $3.8–$6.9 / 100k credits depending on plan. TTS concurrency: 2/5/5/10/15/Custom. STT concurrency: 3/20/20/40/60/Custom. Source: gradium.ai/pricing (July 2026).
05Integrations & Tech Stack
Voice Agent Frameworks
SDKs & API
Deployment
Sovereignty & GDPR
· EU or US residency: paid plan, activation and response signal to verify
· On-device Phonon: separate path; do not infer its guarantees from the cloud API
· ZDR, cloned voices and any on-premise terms: separate contractual scope
· A French company ≠ Paris hosting or automatic Swiss sovereignty
06Relevance for Storygami & Edugami
STT for Storygami & Edugami
Semantic VAD → natural turn detection for AI characters. FR/EN/DE/ES/PT languages covered. Noise robustness in classroom or cinema settings.
Precise TTS + lip-sync
Word-level timestamps and structured-entity pronunciation serve avatar lip-sync and French terminology. Où est Ava ? protagonists’ and Edugami mentor’s emotional prosody still needs testing; no explicit control is documented.
Multilingual S2S Translation
Unique use case: French-speaking student talks, avatar responds in English (or vice versa). Relevant for Edugami in multilingual school contexts.
Real-Time Infrastructure integration
Natively compatible with LiveKit and Pipecat. Aligned with the portal's Real-Time Infrastructure roadmap. Possible evaluation in GamiWays Core Phase B.
Residency & Edugami
Enable the EU endpoint on a paid plan then verify its signal. EU residency is neither Swiss residency nor an automatic guarantee for cloned voice assets; evaluate it against the intended school framework.
07Open Questions
- ?What is the actual TTFA latency of Gradium TTS from Switzerland? (benchmark to perform)
- ?Is Gradium STT WER independently validated (Pipecat benchmark, Artificial Analysis)?
- ?Is EU residency enabled on the selected plan, and does the response signal actually confirm the zone?
- ?Are API pauses and prosody credible enough for Où est Ava ? despite the absence of explicit emotional direction?
- ?Is S2S Translation production-stable (latency, accuracy) for school use in FR→EN?
Related pages
Sources
Data verified: July 2026.