NVIDIA Parakeet TDT 0.6B v3
Open-weights ASR — 25 European languages, timestamps, NVIDIA GPU deployment
Comparative Scores
Architecture
Candidate for a sovereign multi-European ASR track in Phase 2. The published WER and batch RTFx justify a controlled comparison with Audiogami/Whisper on French and Swiss project material; validate 2s chunk behaviour, endpointing, single-stream tail latency and full GPU operations before replacing a streaming provider.
Analysis
Parakeet TDT 0.6B v3 is NVIDIA’s open-weights multilingual ASR model for 25 European languages. Its current model card publishes 6.34% average WER on the Hugging Face Open ASR Leaderboard, including 1.93% on LibriSpeech test-clean, as well as French results on FLEURS, MLS and CoVoST. NVIDIA’s technical report also publishes RTFx 3332.74 batch throughput: on the same protocol, Parakeet’s 6.32% snapshot is ahead of Whisper large-v3’s 7.44% while processing roughly 22.9× faster. NVIDIA documents 2-second chunks for integration, but no conversational TTFA P50/P95; endpointing, tail latency, diarization and GPU cost still need validation on the selected deployment.
Strengths
- 25 European languages including French
- Open weights under CC-BY-4.0
- Word and segment timestamps
- Punctuation and capitalization
- Self-hostable with NeMo / NVIDIA GPU stack
- Long-audio guidance published by NVIDIA
Weaknesses
- No published TTFA P50/P95 or single-stream tail-latency measurement
- Chunked streaming requires integration and endpointing design
- No native diarization documented for the base model
- NVIDIA GPU dependency for the documented production path
- Infrastructure cost must be measured per deployment
STT Capabilities
NVIDIA documents chunked streaming inference for TDT 0.6B v3 through a NeMo script: 2s chunks, 2s right context and 10s left context. This is an integration configuration, not a published TTFA P50/P95 measurement. Endpointing and operations remain implementation responsibilities; Riva/NIM documents a separate Parakeet CTC 0.6B microservice.
25+ languages
Pricing
Model weights are available under CC-BY-4.0. GPU, storage, operations and support costs are infrastructure-dependent and not shown as a universal hourly price.
| Plan | Subscription/mo | Included | Overage/min |
|---|---|---|---|
Self-Hosted — measure infrastructureRecommended Size GPU capacity from a representative load test | Free | GPU / CapEx to measure | — |
Sovereignty & Compliance
Full self-hosting possible with the published weights. Production performance and data residency depend on the selected NVIDIA GPU environment and operations stack.
Data residency: Full control when self-hosted; compliance remains an operator responsibility.
Open weights. Self-host with NeMo or NeMo-Speech.cpp; Riva/NIM documents a separate Parakeet CTC 0.6B production path. NVIDIA GPU sizing and deployment validation remain required.
Open weights. Self-host with NeMo or NeMo-Speech.cpp; Riva/NIM documents a separate Parakeet CTC 0.6B production path. NVIDIA GPU sizing and deployment validation remain required.
This sheet's verification
NVIDIA Parakeet TDT 0.6B v3 model card + technical report (Hugging Face Open ASR Leaderboard)
Update note: 30 septembre 2026 : Moondream Parakeet Redux (178 Mo, CPU) et Ultra (GPU) sont notés, licence CC-BY-4.0, 25 langues. Ils ne remplacent pas cette fiche NVIDIA. Mesures du 19 août conservées : 6,34 % WER moyen, 1,93 % LibriSpeech test-clean, RTFx 3332,74.
This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.
ELO benchmarks / indices: Artificial Analysis