GamiWays
Open SourceCC-BY-4.0Self-hostable

NVIDIA Parakeet TDT 0.6B v3

Open-weights ASR — 25 European languages, timestamps, NVIDIA GPU deployment

3333×
Batch throughput (RTFx)
2s
Streaming chunk (config.)
6.34%
WER (general audio)
Free
Price per minute

Comparative Scores

Accuracy (WER)7/10
Streaming latencyto measure
Multilingual7/10
Sovereignty10/10
Price accessibility9/10
Streaming quality5/10

Architecture

ArchitectureFastConformer encoder + Token-and-Duration Transducer (TDT), NVIDIA NeMo
Parameters0.6B
Languages25+
Self-hostable Yes
Streaming Yes
WER clean audio 1.93%
GamiWays
Phase 2 — Sovereign multilingual ASR

Candidate for a sovereign multi-European ASR track in Phase 2. The published WER and batch RTFx justify a controlled comparison with Audiogami/Whisper on French and Swiss project material; validate 2s chunk behaviour, endpointing, single-stream tail latency and full GPU operations before replacing a streaming provider.

Analysis

Parakeet TDT 0.6B v3 is NVIDIA’s open-weights multilingual ASR model for 25 European languages. Its current model card publishes 6.34% average WER on the Hugging Face Open ASR Leaderboard, including 1.93% on LibriSpeech test-clean, as well as French results on FLEURS, MLS and CoVoST. NVIDIA’s technical report also publishes RTFx 3332.74 batch throughput: on the same protocol, Parakeet’s 6.32% snapshot is ahead of Whisper large-v3’s 7.44% while processing roughly 22.9× faster. NVIDIA documents 2-second chunks for integration, but no conversational TTFA P50/P95; endpointing, tail latency, diarization and GPU cost still need validation on the selected deployment.

Strengths

  • 25 European languages including French
  • Open weights under CC-BY-4.0
  • Word and segment timestamps
  • Punctuation and capitalization
  • Self-hostable with NeMo / NVIDIA GPU stack
  • Long-audio guidance published by NVIDIA

Weaknesses

  • No published TTFA P50/P95 or single-stream tail-latency measurement
  • Chunked streaming requires integration and endpointing design
  • No native diarization documented for the base model
  • NVIDIA GPU dependency for the documented production path
  • Infrastructure cost must be measured per deployment

STT Capabilities

Streaming Yes

NVIDIA documents chunked streaming inference for TDT 0.6B v3 through a NeMo script: 2s chunks, 2s right context and 10s left context. This is an integration configuration, not a published TTFA P50/P95 measurement. Endpointing and operations remain implementation responsibilities; Riva/NIM documents a separate Parakeet CTC 0.6B microservice.

Diarization No
Custom Vocabulary No
Word Timestamps Yes
Auto Punctuation Yes
Multilingual Yes

25+ languages

Pricing

Price / minute
Free
Price / hour
Free
Free tier
Open weights; infrastructure costs excluded

Model weights are available under CC-BY-4.0. GPU, storage, operations and support costs are infrastructure-dependent and not shown as a universal hourly price.

PlanSubscription/moIncludedOverage/min
Self-Hosted — measure infrastructureRecommended

Size GPU capacity from a representative load test

FreeGPU / CapEx to measure—

Sovereignty & Compliance

On-premise Yes

Full self-hosting possible with the published weights. Production performance and data residency depend on the selected NVIDIA GPU environment and operations stack.

GDPR Compliant

Data residency: Full control when self-hosted; compliance remains an operator responsibility.

On-premise Yes

Open weights. Self-host with NeMo or NeMo-Speech.cpp; Riva/NIM documents a separate Parakeet CTC 0.6B production path. NVIDIA GPU sizing and deployment validation remain required.

Self-hosted Deployment

Open weights. Self-host with NeMo or NeMo-Speech.cpp; Riva/NIM documents a separate Parakeet CTC 0.6B production path. NVIDIA GPU sizing and deployment validation remain required.

This sheet's verification

Verified 19 August 2026

NVIDIA Parakeet TDT 0.6B v3 model card + technical report (Hugging Face Open ASR Leaderboard)

Update note: 30 septembre 2026 : Moondream Parakeet Redux (178 Mo, CPU) et Ultra (GPU) sont notés, licence CC-BY-4.0, 25 langues. Ils ne remplacent pas cette fiche NVIDIA. Mesures du 19 août conservées : 6,34 % WER moyen, 1,93 % LibriSpeech test-clean, RTFx 3332,74.

This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.

ELO benchmarks / indices: Artificial Analysis