Inflect-Nano-v2 fine-tuning toolkit released
Key point
A fine-tuning toolkit has been released that lets the ultra-compact TTS model Inflect-Nano-v2 be trained on a specific voice or language.
Details
A new fine-tuning toolkit has been released that lets the ultra-compact TTS models Inflect-Nano-v2 (3.96M parameters) and Inflect-Micro-v2 (9.35M parameters) be trained on a user's own voice or on other languages.
Key features of this toolkit include:
- Custom training: Uses single-speaker recordings and transcripts to warm-start the model or resume training.
- Multilingual support: Handles new phoneme inventories, enabling adaptation to languages other than English.
- Flexible export: Trained results can be exported in PyTorch or ONNX format, allowing deployment across a variety of environments.
- Efficient architecture: A text-to-waveform model that includes a 24kHz waveform decoder, operating with very few parameters.
This toolkit is not zero-shot cloning, but rather a supervised adaptation approach optimized for a specific speaker and language.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.