GamiWays
Cloud APICommercial

Deepgram Aura 2

Aura 2 baseline — Coval: 327ms median TTFA, 5.1% WER (12 Aug 2026)

80 ms
TTFA (best case)
327 ms
TTFA (typical)
$15/1M
Price per million chars
—
ELO Score

Comparative Scores

Voice quality6/10
Latency9/10
Voice cloning1/10
Expressiveness4/10
Performance direction3/10
Sovereignty3/10
Price accessibility7/10
Multilingual1/10

Architecture

ArchitectureProprietary streaming neural TTS
ParametersN/A (cloud)
Languages1
Self-hostable No
Streaming Yes
GamiWays
Phase 1 MVP — Stack ASR+TTS intégré

Relevant for Phase 1 MVP if using Deepgram Nova-3 for ASR. The ASR+TTS stack from a single provider simplifies integration and reduces latency. Limited to English and no voice cloning are significant constraints.

Analysis

Deepgram Aura 2 remains the earlier TTS option in the Deepgram family, optimized for English voice-agent pipelines. The vendor historically claimed <100ms TTFA; the current independent Coval/Openbenchmarks baseline is 327ms median TTFA, 547ms p95 and 5.1% WER for aura-2-thalia-en (12 August 2026). Use it as a measured Deepgram baseline, not as evidence for the new Flux TTS model, whose latency and WER are currently vendor-reported only. No voice cloning, limited expressiveness — focused on predictable English voice-agent operation.

Strengths

  • Independent baseline: 5.1% WER (Coval, 12 Aug 2026)
  • Optimized for voice agent pipelines
  • Natural pairing with Deepgram ASR
  • Reliable at scale

Weaknesses

  • No voice cloning
  • English only (Aura 2)
  • Limited expressiveness
  • No lip-sync data

Voice Capabilities

Voice Cloning No

No voice cloning. Pre-built voices only.

Emotion & performance directionLimited control

Aura provides clean pace and pronunciation control, but no documented programmable emotion.

How: `speed` (0.7–1.5) and IPA pronunciation dictionary.

Validate: For emotional scenes, plan distinct writing and an expressive engine rather than attribute control it does not expose.

Streaming Yes

Vendor claim: <100ms TTFA. Independent Coval/Openbenchmarks measurement (12 Aug 2026): 327ms median TTFA, 547ms p95 and 5.1% WER for Aura-2-Thalia-en. Often paired with Deepgram Nova-3 ASR for a full stack.

Lip-sync Data No

No native lip-sync data.

Pricing

Price / 1M chars
from $15
depending on plan
Price / minute
from $0.0150
depending on plan
Free tier
$200 free credits on signup

$15/1M chars. Often bundled with Deepgram ASR for full voice agent stack.

Hybrid (PAYG + Subscription)
Official pricing page
PlanSubscription/moIncludedOverage/minMax streamsTop-up
Pay As You Go

Same plan as Deepgram STT

Free$200 free credit$0.0150150Yes
GrowthRecommended

Volume discount available

$333/moPrepaid credits, up to 20% discount$0.0120225Yes
Parallel stream limits: The number of simultaneous streams is limited per plan. For events with many concurrent users, verify the required plan or contact the provider for an Enterprise agreement.
Top-up available: Some plans allow purchasing additional minutes (top-up) without changing plans. Useful for occasional usage spikes.

Sovereignty & Compliance

On-premise No

Cloud only. Enterprise on-premise via agreement.

GDPR Compliant

Data residency: US, EU

Strategic & Business Analysis

Deepgram Aura 2 — Strategic Positioning

Beyond technical specs: where does this tool sit in the ecosystem, what are the risks and strategic implications for GamiWays?

Deepgram Aura is the unified voice platform play: STT+TTS+LLM in one API, sub-200ms latency, $1.3B valuation — but its VPC-only stance limits sovereignty appeal for regulated European deployments.

Cloud + VPC
Lock-in risk:Medium
Sovereignty fit:Medium
Open-source threat:Medium
Pricing:Commoditizing ↓↓

A. Strategic Positioning

Target customer: Developer / Enterprise — real-time voice agents, unified STT+TTS platform

Sub-200ms TTS latency as part of a unified STT+TTS+LLM platform — the one-stop shop for voice AI agent infrastructure.

B. Competitive Moat

  • Unified STT+TTS+LLM API — reduces integration complexity and latency vs multi-vendor stacks
  • Sub-200ms TTS latency with natural voices — competitive with Cartesia for real-time agents
  • Series C $130M (Jan 2026) — $1.3B valuation — financial strength for R&D

Vulnerability: No on-premise TTS option (VPC only). Open-source models catching up. Pricing pressure from competitors.

E. Strategic Questions for GamiWays

Sovereignty fit

EU data residency via VPC available. No full on-premise for TTS. Moderate sovereignty fit — better than pure cloud, worse than Inworld.

Build vs. Buy

Buy for Phase 1 (unified platform, low integration overhead). Evaluate on-premise alternatives for Phase 2 sovereignty.

Lock-in risk

Unified STT+TTS platform creates integration lock-in. VPC deployment and competitive pricing reduce dependency risk.

Roadmap alignment

Good for Phase 1 (unified STT+TTS for voice agents). Phase 2 requires on-premise TTS — consider Inworld or open-source for sovereignty.

This sheet's verification

Verified 12 August 2026

Coval/Openbenchmarks — Aura-2-Thalia-en, 12 Aug 2026 (327ms median TTFA, 547ms p95, 5.1% WER)

Update note: Mise à jour Flux TTS : baseline Aura-2 indépendante ajoutée depuis Coval/Openbenchmarks. Les mesures Aura ne sont pas attribuables à Flux TTS.

This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.