STT / Speech-to-text
Comparison of speech recognition solutions for conversational pipelines (2025–2026). Benchmarks, strategic stakes, and decision questions.
Strategic framing: read STT in context
What the numbers say
WER describes accuracy on a given corpus; it does not predict proper nouns, accents, noise or turn-taking in your setting. TTFA describes stream start, not conversational responsiveness on its own.
What to combine in the table
For live interaction, read WER and TTFA with streaming, diarization and languages. Endpointing, turn detection and quality on your recordings often matter more than the best isolated score.
Stay alert to
Data residency and retention, stream limits, PII redaction and total cost. The hourly price compares the first public tier; it does not replace load testing or a cloud/self-hosted fallback plan.
Reading method: filter non-negotiable requirements first, then compare accuracy and real-time performance on the same use case, and finally validate with representative audio, languages and noise conditions.
How are the table metrics measured?open
WER: error rate on a specific corpus. A value is comparable only with the same corpus, language, model variant and protocol; it does not directly predict performance on your audio.
TTFA: time to the first transcription fragment for a given stream. “—” means no traceable P50/P95 measurement is published, not zero latency.
RTFx: offline or batch processing throughput: 100× processes one hour of audio in about 36 seconds. It is neither conversational latency nor a single-stream real-time promise.
Streaming chunks: integration settings that define the audio portion and context sent to the model. They must not be read as TTFA. Example: Parakeet’s 2s value is a NeMo chunk, not a latency measurement.
Method: each row gives its source and verification date in the full record. For a decision, compare values on a common protocol first, then test the project’s audio and infrastructure.
Real-time cloud STT APIs — Deepgram Nova-3 is the latency reference (75ms TTFA). Whisper large-v3 is the open-source quality standard. AssemblyAI Universal-3 Pro leads the multilingual WER benchmark — Voice Agent API $4.50/hr (launched Apr 29, 2026).
Click a header to sort
| Solution | TTFA | WER | Streaming | Multilingual | Diarization | Price/hr · entry | Sovereign |
|---|---|---|---|---|---|---|---|
Google Gemini 3.5 TranscribeNew Live STT + recorded-audio analysis — benchmark candidate | — | — | ✓ | 85 langs | ✗ | $0.54/h | ✗ |
Deepgram Nova-3Upd Phase 1 MVP — Streaming ASR | 75ms | ✓ | 45 langs | ✓ | $0.26/h | ✓ | |
Inworld STTUpd Phase 1 MVP — Emotional STT + Avatar behaviour Axis 2 | 92ms | ✓ | 100 langs | ✓ | $0.10/h | ✗ | |
Gradium STTNew Phase B R&D — Turn detection + real-time infrastructure | 100ms | ✓ | 5 langs | ✓ | $0.62/h | ✓ | |
AssemblyAI Universal-3.6 Pro RealtimeUpd Voice Agent Pipeline — Accuracy reference | 150ms | ✓ | 32 langs | ✓ | $0.21/h | ✗ | |
Azure Speech (Microsoft) Swiss sovereignty — institutional | 180ms | ✓ | 100 langs | ✓ | $1.00/h | ✓ | |
xAI Grok STTNew Phase B R&D — S2S pipeline evaluation | 200ms | ✓ | 25 langs | ✓ | $0.10/h | ✗ | |
Soniox v5New Accuracy + cost — production validation needed | 249ms | ✓ | 50 langs | ✓ | $0.12/h | ✗ | |
Cartesia Ink-2New Phase 1 MVP — Full-stack Cartesia pipeline (STT + TTS) | 299ms | ✓ | 40 langs | ✗ | $0.54/h | ✗ |
Compare by solution
Google Gemini 3.5 Transcribe separates real-time speech recognition from recorded-audio transcription. Gemini 3.5 Transcribe Live streams interim and final text over the Live API and offers automatic, hybrid or manual VAD plus custom vocabulary. The file model adds speaker diarization and word timestamps, but those two features are not available in the live path. Google announces model-specific WER figures, yet they use a different protocol from the portal’s benchmarks; comparative accuracy and end-to-end latency remain to be measured. Both paths are Gemini Developer API cloud services, not Google Cloud Speech-to-Text v2/Chirp.
Full details →Deepgram Nova-3 is a real-time ASR option for voice agents, with 45+ language support, built-in VAD and endpointing. The Voice Agent API provides a complete STT+LLM+TTS pipeline in a single WebSocket, with configurable STT, LLM providers, TTS and function calling for RAG integration. Its current standard pay-as-you-go rate is promotional through September 12, 2026, then reverts to $0.075/min. Deepgram's TTS family includes Aura-2 and Flux TTS; external TTS can still be integrated. Enterprise on-premise deployment and an EU endpoint are available. The 75ms P90 value remains a historical Pipecat benchmark, not a live Deepgram SLA.
Full details →Inworld STT (2025–2026) is the most feature-rich cloud STT API for interactive voice agents. Sub-100ms documented latency, 100+ languages via multi-provider routing (Whisper large-v3 + AssemblyAI). Unique real-time voice profiling extracts emotion (happy/calm/angry/frustrated), accent, age, pitch, and vocal style on every streaming chunk. Realtime API (full pipeline STT+LLM+TTS) from $0.015/min — 4x cheaper than OpenAI Realtime ($0.06/min). Native voice cloning: built-in + cloned + custom voices (up to 3,000 custom voices on Growth plan). RAG via function calling (tool calling mid-conversation). ZDR support. On-premise available on Enterprise. Drop-in compatible with OpenAI Realtime API.
Full details →Gradium STT features semantic VAD for intelligent turn detection (meaning-based, not just silence). Best-in-class accuracy claimed by founders who invented neural audio codecs. 5 languages (FR/EN/DE/ES/PT), word-level timestamps, speaker diarization. Native LiveKit and Pipecat integration. On-premise Enterprise option with zero data retention.
Full details →On 29 September 2026 AssemblyAI released Universal-3.6 Pro Realtime (`universal-3-6-pro`): 32 languages, $0.45/hr, and entity-aware endpointing. Its English voice-agent WER of 5.19% is a vendor protocol and is not the sheet’s comparative WER. Earlier context: AssemblyAI launched its Voice Agent API on April 29, 2026: a complete STT+LLM+TTS pipeline in one WebSocket connection at a flat $4.50/hr. Universal-3 Pro Streaming (u3-rt-pro) is its real-time STT model, with semantic and acoustic turn detection, native barge-in, and session resumption. JSON Schema tool calling supports a custom RAG stack (Pinecone, LlamaIndex, and others) through function calling; there is no native RAG, but external integration is complete. There is no native voice cloning: voices are predefined (18+ English and multilingual voices), though an external TTS such as ElevenLabs or Cartesia can provide a custom voice. LeMUR features cover transcription summarisation, Q&A and sentiment. Supports 99 languages, diarisation and word-level timestamps.
Full details →Azure Speech offers enterprise-grade STT with 100+ languages, custom model training, and Swiss data center (Zurich). 5.9% WER on English. Disconnected container deployment for partial sovereignty. Most expensive option but strongest enterprise compliance. Swiss German custom model available.
Full details →Grok STT is the standalone transcription API built on the same stack as Grok Voice Think Fast 2.0. It achieves 6.9% WER globally across phone calls, meetings, video/podcasts and telephony — outperforming ElevenLabs Scribe v2 (9.0%), Deepgram Nova-3 (11.0%) and AssemblyAI (12.9%) in xAI's internal benchmark. Particularly strong on entity recognition (phone numbers, dates, currencies) thanks to advanced Inverse Text Normalization. Supports 25+ languages, word-level timestamps, speaker diarization, multichannel audio, and keyterm biasing. Priced at $0.10/h batch — among the most competitive in the market. Best evaluated as part of the Grok Voice Think Fast 2.0 S2S pipeline ($0.08/min all-in).
Full details →Soniox v5 claims the best WER on the market at 1.25% (internal benchmark). Priced at $0.12/hr — the most affordable STT API in the benchmark. 249ms TTFS, 50+ languages, speaker diarization. On-premise available on Enterprise. Very new — production track record limited. New entry June 2026.
Full details →Cartesia Ink-2 achieves the best WER on the Pipecat STT Benchmark (June 2026) at 1.47%, ahead of Deepgram (1.71%) and AssemblyAI (1.74%). Priced at $0.43/hr with 299ms TTFS. As of 30 September 2026 Cartesia still describes this voice-agent model as English-only, with multilingual support announced and not shipped. The paired TTS on this portal is now Sonic 3.6, which does speak French. Ink-2 is therefore not a French STT for Où est Ava ?.
Full details →