AI Briefing
KOSign in

Sopro V2 Turbo 2610 Update Improves Cloned Voice Quality on 120M CPU Model

·2026.10.04 01:48

Key point

The update reduces voice roughness and break-up while maintaining ~300 ms first-audio latency on laptop CPUs.

Details

Sopro V2 Turbo 2610 is an interim update to the 120M parameter text-to-speech model, focusing on reducing roughness and break-up in cloned voices. The model maintains its lightweight architecture, delivering approximately 300 ms time-to-first-audio on a standard laptop CPU.

Supported Languages and Performance

The update supports English, European Portuguese, French, and German, with additional languages planned for future releases. The model remains Apache-2.0 licensed and is designed for true streaming and local execution without GPU requirements.

Known Limitations

Current limitations include struggles with very high-pitched or cartoon-like voices, noisy reference audio, and some out-of-distribution (OOD) voices. The developers are actively seeking failed samples to improve these edge cases.

Availability

The model is available via uvx --from sopro soprotts serve, with weights hosted on Hugging Face and an in-browser demo for desktop users.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.