AI Briefing
KO

NVIDIA Voice Stack Localized

·2026.08.07 07:54

Key point

NVIDIA supports on-device execution by quantizing ASR, TTS, and speech codecs to GGUF.

Details

Speech recognition (ASR), text-to-speech (TTS), and speech codecs can now be run in local environments via NVIDIA's NeMo-Speech.cpp.

The supported models and components are as follows:

  • Magpie-TTS Multilingual
  • Nemotron Speech Streaming EN 0.6B
  • Nemotron-3.5 ASR Streaming
  • Parakeet CTC 1.1B
  • Parakeet TDT 0.6B v3
  • NanoCodec

The models are quantized in GGUF format, and Hugging Face model cards provide instructions for local execution using NeMo-Speech.cpp. NanoCodec support has also been merged into the project.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.