Google Gemini 3.5 Transcribe
Live API STT — streaming bidirectionnel, 85+ locales, VAD hybride et vocabulaire personnalisé
Comparative Scores
Architecture
Candidate for a controlled Où est Ava ? streaming-STT comparison when hybrid VAD, custom French vocabulary and interim text can improve turn timing. Test the 10-minute session boundary, finalization latency, interruptions and French character names before using it in a live experience. For Plastic Dilemma recordings, compare the file path’s diarization/timestamps with a verbatim transcript rather than assuming the live model covers the same features.
Analysis
Google Gemini 3.5 Transcribe separates real-time speech recognition from recorded-audio transcription. Gemini 3.5 Transcribe Live streams interim and final text over the Live API and offers automatic, hybrid or manual VAD plus custom vocabulary. The file model adds speaker diarization and word timestamps, but those two features are not available in the live path. Google announces model-specific WER figures, yet they use a different protocol from the portal’s benchmarks; comparative accuracy and end-to-end latency remain to be measured. Both paths are Gemini Developer API cloud services, not Google Cloud Speech-to-Text v2/Chirp.
Strengths
- Dedicated bidirectional Live API STT with interim and final transcripts
- 85+ locales with automatic language detection and code-switching
- Hybrid VAD can finalize a turn from a client-side silence signal
- Up to 1,000 custom vocabulary terms
- Recorded-audio path adds diarization and word-level timestamps
- Live API ecosystem integrations: LiveKit, Pipecat, Agora and Vercel
Weaknesses
- No published TTFA P50/P95 or end-to-end dialog latency
- Live sessions limited to 10 minutes
- No live diarization or word-level timestamps
- Smart transcription is incompatible with diarization and word timestamps
- Cloud only; Gemini API data residency and retention must be assessed separately
- Google-reported WER not comparable with the portal’s common benchmark
STT Capabilities
Live API WebSocket : textes intermédiaires et finalisés, détection automatique ou indice BCP-47, vocabulaire personnalisé et VAD automatique/hybride/manuelle. Session continue limitée à 10 min ; aucune TTFA P50/P95 traçable n’est publiée.
85+ languages
Pricing
Gemini 3.5 Transcribe Live : ~0,009$/min, soit ~0,54$/h à partir des taux audio+texte estimés par Google. Fichiers (Interactions API) : ~0,005$/min, soit ~0,30$/h. Les deux tarifs Developer API sont des estimations basées sur les tokens ; ils ne sont pas les tarifs Cloud STT v2/Chirp.
Sovereignty & Compliance
Cloud API only. Operator must assess the Google data path, retention and compliance terms for the chosen account and region.
Data residency: No regional residency claim is carried over from Cloud STT v2; verify the Gemini API terms and deployment path for the intended use.
Gemini Developer API cloud only. No on-premise option documented for these models.
Google Gemini 3.5 Transcribe — Strategic Positioning
Beyond technical specs: where does this tool sit in the ecosystem, what are the risks and strategic implications for GamiWays?
Gemini 3.5 Transcribe adds a credible Live API STT comparison path, particularly where hybrid VAD and custom vocabulary influence turn timing — but Google-reported WER and total dialog latency must not be treated as common-benchmark evidence, and the service remains cloud-only.
A. Strategic Positioning
Target customer: Developers building real-time and recorded-audio workflows in the Gemini API ecosystem
Gemini 3.5 Transcribe separates low-latency live transcription from recorded-audio transcription, adding hybrid VAD, custom vocabulary, diarization and word timestamps by path.
B. Competitive Moat
- Dedicated Live API STT with interim and final transcripts over a bidirectional connection
- Hybrid VAD and custom vocabulary provide control points for conversational integration
- Recorded-audio path adds speaker diarization and word-level timestamps
Vulnerability: Vendor lock-in risk with Google Cloud. Open-source models catching up. No on-premise option outside specific partnerships.
E. Strategic Questions for GamiWays
Sovereignty fit
EU continental boundary available but cloud-only. Google Cloud dependency creates sovereignty risk for Swiss/EU regulated deployments.
Build vs. Buy
Buy for Phase 1 multilingual requirements. For Phase 2 sovereignty, switch to Whisper/Voxtral self-hosted to eliminate Google dependency.
Lock-in risk
Deep Google Cloud ecosystem integration creates strong lock-in. Switching costs are high if Vertex AI or Gemini are also used.
Roadmap alignment
Good for Phase 1 multilingual transcription. Incompatible with Phase 2 sovereignty requirements without major architectural changes.
This sheet's verification
Google Gemini 3.5 Transcribe announcement and API documentation — common-protocol benchmark pending
Update note: Vérifié le 31 août 2026 : Gemini 3.5 Transcribe Live ($0,005/min audio + $0,004/min texte, ~0,009$/min) et Gemini 3.5 Transcribe fichiers ($0,003 + $0,002/min, ~0,005$/min). Les WER Google 4,0% live / 2,6% fichiers sont cités dans la fiche mais exclus du score comparatif, faute de protocole commun. Limites live : 10 min, sans diarisation ni timestamps mot-à-mot.
This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.
ELO benchmarks / indices: Artificial Analysis