NVIDIA Voice Stack Localized
·2026.08.07 07:54
Key point
NVIDIA supports on-device execution by quantizing ASR, TTS, and speech codecs to GGUF.
Details
Speech recognition (ASR), text-to-speech (TTS), and speech codecs can now be run in local environments via NVIDIA's NeMo-Speech.cpp.
The supported models and components are as follows:
- Magpie-TTS Multilingual
- Nemotron Speech Streaming EN 0.6B
- Nemotron-3.5 ASR Streaming
- Parakeet CTC 1.1B
- Parakeet TDT 0.6B v3
- NanoCodec
The models are quantized in GGUF format, and Hugging Face model cards provide instructions for local execution using NeMo-Speech.cpp. NanoCodec support has also been merged into the project.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.