Deepgram Nova-3
Real-time ASR — 45+ languages, EU endpoint, Enterprise on-premise option
Comparative Scores
Architecture
Primary candidate for Phase 1 MVP ASR. 75ms latency is critical for sub-2s pipeline. On-premise option aligns with Swiss sovereignty requirements. Audiogami (Gamilab) already in production — Deepgram as fallback/comparison.
Analysis
Deepgram Nova-3 is a real-time ASR option for voice agents, with 45+ language support, built-in VAD and endpointing. The Voice Agent API provides a complete STT+LLM+TTS pipeline in a single WebSocket, with configurable STT, LLM providers, TTS and function calling for RAG integration. Its current standard pay-as-you-go rate is promotional through September 12, 2026, then reverts to $0.075/min. Deepgram's TTS family includes Aura-2 and Flux TTS; external TTS can still be integrated. Enterprise on-premise deployment and an EU endpoint are available. The 75ms P90 value remains a historical Pipecat benchmark, not a live Deepgram SLA.
Strengths
- Voice Agent API $3.36/hr through Sept. 12, then $4.50/hr standard — configurable pipeline + function calling
- Built-in VAD + endpointing
- 45+ languages
- Enterprise on-premise option
- EU endpoint available
- LiveKit/Pipecat native integration
- Speaker diarization + word timestamps
Weaknesses
- Cloud-first (sovereignty limited — on-premise Enterprise only)
- AA-WER 5.2% is not among the accuracy leaders
- French/German WER must be tested on project material
- No native voice cloning — Aura-2 preset voices and Flux cloning not available at launch
- No native RAG — requires external function calling infrastructure
- No open-weights
STT Capabilities
WebSocket streaming. 75ms P90 latency. Interim results with endpointing. Voice Activity Detection (VAD) built-in.
45+ languages
Pricing
$0.26/hr ($0.0043/min) Nova-3 Monolingual Pay As You Go, vérifié le 18 août 2026. Growth : $0.0036/min. Le benchmark Pipecat historique reste cité séparément ; on-premise : prix Enterprise sur devis.
| Plan | Subscription/mo | Included | Overage/min | Max streams | Top-up |
|---|---|---|---|---|---|
Pay As You Go Pre-pay credits at PAYG rate | Free | $200 free credit | $0.0043 | 150 | Yes |
GrowthRecommended Additional prepaid credits at Growth rate | $333/mo | Prepaid credits, up to 20% discount | $0.0036 | 225 | Yes |
Enterprise Custom volume pricing | Free | Custom | — | ∞ | Yes |
Sovereignty & Compliance
Self-hosted Docker deployment. Enterprise on-premise available. Partial sovereignty.
Data residency: US (default). EU data residency available.
On-premise deployment available (enterprise). Self-hosted via Docker.
On-premise deployment available (enterprise). Self-hosted via Docker.
Deepgram Nova-3 — Strategic Positioning
Beyond technical specs: where does this tool sit in the ecosystem, what are the risks and strategic implications for GamiWays?
Deepgram Nova-3 is the enterprise STT leader with 54.2% lower WER on noisy audio — but its VPC-only stance (no full on-premise) limits sovereignty appeal for regulated European deployments.
A. Strategic Positioning
Target customer: Developer / Enterprise — voice agents, real-time transcription
Unified STT+TTS+LLM platform with 54.2% lower WER on noisy audio vs competitors — the voice AI infrastructure backbone.
B. Competitive Moat
- 54.2% lower WER on noisy audio vs competitors including hyperscalers
- Unified STT+TTS+LLM API — reduces integration complexity and end-to-end latency
- Series C $130M (Jan 2026) — $1.3B valuation — financial strength for R&D
Vulnerability: Open-source models (Whisper, Voxtral) catching up. No full on-premise STT option (VPC only). Pricing pressure from competitors.
E. Strategic Questions for GamiWays
Sovereignty fit
EU data residency via VPC available. No full on-premise STT. Moderate sovereignty fit — better than pure cloud, worse than self-hosted.
Build vs. Buy
Buy for Phase 1 (best accuracy, unified platform). Evaluate Whisper/Voxtral self-hosted for Phase 2 sovereignty.
Lock-in risk
Unified STT+TTS platform creates integration lock-in. VPC deployment and competitive pricing reduce dependency risk.
Roadmap alignment
Good for Phase 1 voice agents. Phase 2 sovereignty requires on-premise STT — consider Whisper or Voxtral self-hosted.
This sheet's verification
Inworld benchmark 2026 + Koenecke et al.
Update note: Vérification 18 août 2026 : AA-WER v2 5,2% et facteur de vitesse 541,1× via Artificial Analysis ; Nova-3 Monolingual Pay As You Go $0,0043/min via Deepgram. Voice Agent API Standard : $0,056/min jusqu’au 12 septembre, puis $0,075/min. Les mesures Pipecat de juin (1,71% WER, 75ms P90) sont conservées comme benchmark historique distinct, non comme mesures API actuelles.
This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.
ELO benchmarks / indices: Artificial Analysis