GamiWays
Cloud APICommercial

Google Gemini 3.5 Transcribe

Live API STT — streaming bidirectionnel, 85+ locales, VAD hybride et vocabulaire personnalisé

—
Latency (best case)
—
Latency (typical)
—
WER (general audio)
$0.0090/min
Price per minute

Comparative Scores

Accuracy (WER)to measure
Streaming latencyto measure
Multilingual9/10
Sovereignty1/10
Price accessibility7/10
Streaming quality8/10

Architecture

ArchitectureGemini 3.5 Transcribe Live via Live API (WebSocket) ; Gemini 3.5 Transcribe fichiers via Interactions API
ParametersN/A (cloud)
Languages85+
Self-hostable No
Streaming Yes
WER clean audio 0%
GamiWays
Live STT + recorded-audio analysis — benchmark candidate

Candidate for a controlled Où est Ava ? streaming-STT comparison when hybrid VAD, custom French vocabulary and interim text can improve turn timing. Test the 10-minute session boundary, finalization latency, interruptions and French character names before using it in a live experience. For Plastic Dilemma recordings, compare the file path’s diarization/timestamps with a verbatim transcript rather than assuming the live model covers the same features.

Analysis

Google Gemini 3.5 Transcribe separates real-time speech recognition from recorded-audio transcription. Gemini 3.5 Transcribe Live streams interim and final text over the Live API and offers automatic, hybrid or manual VAD plus custom vocabulary. The file model adds speaker diarization and word timestamps, but those two features are not available in the live path. Google announces model-specific WER figures, yet they use a different protocol from the portal’s benchmarks; comparative accuracy and end-to-end latency remain to be measured. Both paths are Gemini Developer API cloud services, not Google Cloud Speech-to-Text v2/Chirp.

Strengths

  • Dedicated bidirectional Live API STT with interim and final transcripts
  • 85+ locales with automatic language detection and code-switching
  • Hybrid VAD can finalize a turn from a client-side silence signal
  • Up to 1,000 custom vocabulary terms
  • Recorded-audio path adds diarization and word-level timestamps
  • Live API ecosystem integrations: LiveKit, Pipecat, Agora and Vercel

Weaknesses

  • No published TTFA P50/P95 or end-to-end dialog latency
  • Live sessions limited to 10 minutes
  • No live diarization or word-level timestamps
  • Smart transcription is incompatible with diarization and word timestamps
  • Cloud only; Gemini API data residency and retention must be assessed separately
  • Google-reported WER not comparable with the portal’s common benchmark

STT Capabilities

Streaming Yes

Live API WebSocket : textes intermédiaires et finalisés, détection automatique ou indice BCP-47, vocabulaire personnalisé et VAD automatique/hybride/manuelle. Session continue limitée à 10 min ; aucune TTFA P50/P95 traçable n’est publiée.

Diarization No
Custom Vocabulary Yes
Word Timestamps No
Auto Punctuation Yes
Multilingual Yes

85+ languages

Pricing

Price / minute
from $0.0090
depending on plan
Price / hour
from $0.540
depending on plan
Free tier
Free tier available; limits vary by API and account

Gemini 3.5 Transcribe Live : ~0,009$/min, soit ~0,54$/h à partir des taux audio+texte estimés par Google. Fichiers (Interactions API) : ~0,005$/min, soit ~0,30$/h. Les deux tarifs Developer API sont des estimations basées sur les tokens ; ils ne sont pas les tarifs Cloud STT v2/Chirp.

Sovereignty & Compliance

On-premise No

Cloud API only. Operator must assess the Google data path, retention and compliance terms for the chosen account and region.

GDPR No

Data residency: No regional residency claim is carried over from Cloud STT v2; verify the Gemini API terms and deployment path for the intended use.

On-premise No

Gemini Developer API cloud only. No on-premise option documented for these models.

Strategic & Business Analysis

Google Gemini 3.5 Transcribe — Strategic Positioning

Beyond technical specs: where does this tool sit in the ecosystem, what are the risks and strategic implications for GamiWays?

Gemini 3.5 Transcribe adds a credible Live API STT comparison path, particularly where hybrid VAD and custom vocabulary influence turn timing — but Google-reported WER and total dialog latency must not be treated as common-benchmark evidence, and the service remains cloud-only.

Cloud SaaS only
Lock-in risk:High
Sovereignty fit:Low
Open-source threat:High
Pricing:Falling ↓

A. Strategic Positioning

Target customer: Developers building real-time and recorded-audio workflows in the Gemini API ecosystem

Gemini 3.5 Transcribe separates low-latency live transcription from recorded-audio transcription, adding hybrid VAD, custom vocabulary, diarization and word timestamps by path.

B. Competitive Moat

  • Dedicated Live API STT with interim and final transcripts over a bidirectional connection
  • Hybrid VAD and custom vocabulary provide control points for conversational integration
  • Recorded-audio path adds speaker diarization and word-level timestamps

Vulnerability: Vendor lock-in risk with Google Cloud. Open-source models catching up. No on-premise option outside specific partnerships.

E. Strategic Questions for GamiWays

Sovereignty fit

EU continental boundary available but cloud-only. Google Cloud dependency creates sovereignty risk for Swiss/EU regulated deployments.

Build vs. Buy

Buy for Phase 1 multilingual requirements. For Phase 2 sovereignty, switch to Whisper/Voxtral self-hosted to eliminate Google dependency.

Lock-in risk

Deep Google Cloud ecosystem integration creates strong lock-in. Switching costs are high if Vertex AI or Gemini are also used.

Roadmap alignment

Good for Phase 1 multilingual transcription. Incompatible with Phase 2 sovereignty requirements without major architectural changes.

This sheet's verification

Verified 31 August 2026

Google Gemini 3.5 Transcribe announcement and API documentation — common-protocol benchmark pending

Update note: Vérifié le 31 août 2026 : Gemini 3.5 Transcribe Live ($0,005/min audio + $0,004/min texte, ~0,009$/min) et Gemini 3.5 Transcribe fichiers ($0,003 + $0,002/min, ~0,005$/min). Les WER Google 4,0% live / 2,6% fichiers sont cités dans la fiche mais exclus du score comparatif, faute de protocole commun. Limites live : 10 min, sans diarisation ni timestamps mot-à-mot.

This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.

ELO benchmarks / indices: Artificial Analysis