GamiWays
Cloud APICommercial

AssemblyAI Universal-3.6 Pro Realtime

universal-3-6-pro (29 sept. 2026) — 32 langues, 0,45 $/h, endpointing d’entités. WER fiche inchangé.

150ms
Latency (best case)
300ms
Latency (typical)
3.1%
WER (general audio)
$0.0075/min
Price per minute

Comparative Scores

Accuracy (WER)10/10
Streaming latency7/10
Multilingual10/10
Sovereignty1/10
Price accessibility5/10
Streaming quality9/10

Architecture

ArchitectureUniversal-3 Pro (prompt-based transformer, domain customization sans retraining). Universal-3 Pro Streaming (u3-rt-pro) pour voice agents temps réel. Voice Agent API = WebSocket unique STT+LLM+TTS.
ParametersN/A (cloud)
Languages32+
Self-hostable No
Streaming Yes
WER clean audio 1.1%
GamiWays
Voice Agent Pipeline — Accuracy reference

Voice Agent API très pertinent pour GamiWays Phase 1 : pipeline STT+LLM+TTS en 1 WebSocket simplifie l'architecture. Tool calling permet d'intégrer un RAG sur la base de connaissances GamiWays. Pas de clonage vocal natif : intégrer ElevenLabs ou Cartesia via TTS externe pour la voix du GamiWays. Référence précision WER pour benchmarking.

Analysis

On 29 September 2026 AssemblyAI released Universal-3.6 Pro Realtime (`universal-3-6-pro`): 32 languages, $0.45/hr, and entity-aware endpointing. Its English voice-agent WER of 5.19% is a vendor protocol and is not the sheet’s comparative WER. Earlier context: AssemblyAI launched its Voice Agent API on April 29, 2026: a complete STT+LLM+TTS pipeline in one WebSocket connection at a flat $4.50/hr. Universal-3 Pro Streaming (u3-rt-pro) is its real-time STT model, with semantic and acoustic turn detection, native barge-in, and session resumption. JSON Schema tool calling supports a custom RAG stack (Pinecone, LlamaIndex, and others) through function calling; there is no native RAG, but external integration is complete. There is no native voice cloning: voices are predefined (18+ English and multilingual voices), though an external TTS such as ElevenLabs or Cartesia can provide a custom voice. LeMUR features cover transcription summarisation, Q&A and sentiment. Supports 99 languages, diarisation and word-level timestamps.

Strengths

  • Voice Agent API : pipeline STT+LLM+TTS complet en 1 WebSocket, $4.50/hr flat
  • Universal-3 Pro Streaming : turn detection sémantique+acoustique, barge-in natif
  • Tool calling JSON Schema → RAG custom intégrable (Pinecone, LlamaIndex, etc.)
  • Session resumption 30s, live config update mid-conversation
  • 4.9% WER Universal-2 — meilleure précision cloud
  • 99 langues, diarisation, LeMUR AI (résumé, Q&R, sentiment)
  • GDPR, SOC 2 Type 2, ISO 27001, HIPAA, PCI DSS

Weaknesses

  • Pas de clonage vocal natif — voix prédéfinies uniquement (TTS externe possible via intégration)
  • Pas de RAG natif — nécessite tool calling + infrastructure RAG externe
  • Cloud only — pas d'option on-premise, souveraineté limitée
  • Voice Agent API : LLM intégré non personnalisable nativement (custom LLM via Retell AI possible)
  • $4.50/hr Voice Agent API — coût élevé pour usage intensif

STT Capabilities

Streaming Yes

universal-3-6-pro : même endpoint streaming, nom de modèle à changer. Endpointing qui garde le tour ouvert sur une entité. Latence médiane d’endpoint 537 ms, chiffre AssemblyAI, identique à 3.5 Pro.

Diarization Yes
Custom Vocabulary Yes
Word Timestamps Yes
Auto Punctuation Yes
Multilingual Yes

32+ languages

Pricing

Price / minute
from $0.0075
depending on plan
Price / hour
from $0.450
depending on plan
Free tier
$50 credit on signup — no credit card required

universal-3-6-pro : 0,45 $/h selon AssemblyAI le 29 septembre 2026. Le 0,57 $/h Pipecat de juin et le Voice Agent à 4,50 $/h restent des repères antérieurs.

PlanSubscription/moIncludedOverage/minTop-up
Free

5 new streams/min rate limit on free tier

Free$50 free credit$0.0025—
Pay As You GoRecommended

Unlimited concurrency — no stream cap documented

FreeNone$0.0025Yes
Top-up available: Some plans allow purchasing additional minutes (top-up) without changing plans. Useful for occasional usage spikes.

Sovereignty & Compliance

On-premise No

Cloud only. No on-premise. EU data residency available (GDPR, SOC 2 Type 2, ISO 27001, HIPAA, PCI DSS).

GDPR Compliant

Data residency: US (default). EU data residency disponible. SOC 2 Type 2, ISO 27001, HIPAA, PCI DSS.

On-premise No

Cloud only. No on-premise option. EU data residency available.

Strategic & Business Analysis

AssemblyAI Universal-3.6 Pro Realtime — Strategic Positioning

Beyond technical specs: where does this tool sit in the ecosystem, what are the risks and strategic implications for GamiWays?

AssemblyAI is the audio intelligence leader — #1 accuracy on Hugging Face leaderboard, 30% fewer hallucinations, full PII/diarization suite. EU Dublin data residency available but no on-premise limits Phase 2 sovereignty.

Cloud + VPC
Lock-in risk:Medium
Sovereignty fit:Medium
Open-source threat:Medium
Pricing:Commoditizing ↓↓

A. Strategic Positioning

Target customer: Developer / Enterprise — audio intelligence, voice agents, Fortune 500

Ranked #1 on Hugging Face Open ASR Leaderboard with Universal-3 Pro — 30% fewer hallucinations than competitors, full audio intelligence suite.

B. Competitive Moat

  • #1 on Hugging Face Open ASR Leaderboard — Universal-3 Pro with 30% fewer hallucinations
  • Full audio intelligence suite: diarization, PII redaction, content moderation — beyond transcription
  • SOC 2 Type 2, PCI-DSS 4.0 Level 1, ISO 27001 in progress — enterprise compliance

Vulnerability: Open-source Whisper catching up in quality. High switching costs if deeply integrated. No full on-premise option.

E. Strategic Questions for GamiWays

Sovereignty fit

EU data residency in Dublin available. No full on-premise. Strong compliance certifications reduce regulatory risk.

Build vs. Buy

Buy for Phase 1 (best accuracy, audio intelligence suite). For Phase 2 sovereignty, evaluate Whisper self-hosted for basic transcription.

Lock-in risk

Developer-focused API creates integration dependency. Audio intelligence suite features increase switching costs.

Roadmap alignment

Good for Phase 1 voice agents and audio intelligence. Phase 2 sovereignty requires self-hosted alternatives for full data control.

This sheet's verification

Verified 30 September 2026

AssemblyAI Voice Agent API launch (Apr 29, 2026) + pricing page

Update note: Universal-3.6 Pro Realtime vérifié le 30 septembre 2026 : modèle universal-3-6-pro, 32 langues, 0,45 $/h, endpoint 537 ms (éditeur). Le WER AA/Pipecat de la fiche n’est pas écrasé.

This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.

ELO benchmarks / indices: Artificial Analysis