AI Briefing
KO

Quantization Backfires

·2026.04.15 23:01

Key point

After testing 30 TTS engines on a MacBook Air M4, Kokoro turned out to be the optimal choice, and quantization actually made things slower.

Details

In the real-time translation pipeline, TTS was the bottleneck. Deepgram Nova-3 STT took about 300ms and Groq Llama 3.3 70B translation took about 200ms, adding up to around 500ms, but once TTS exceeded 1 second, the sense of conversation fell apart.

Comparing local TTS on the same sentences on a MacBook Air M4, 24GB RAM:

  • Piper ryan-med: 30~50ms (2~3 words), 137ms (10 words) — fastest, but quality was B
  • Kokoro 82M fp16: 370ms (2~3 words), 730ms (10 words), quality A+ — the best overall balance
  • pocket-tts: 260ms / 7500ms — impractical beyond short sentences
  • ZipVoice 123M: about 500ms / 1240ms, quality B+
  • Chatterbox 500M: 6310ms / 9100ms, quality A but too slow for real-time use
  • Qwen3-TTS 0.6B: about 800ms / 1600~2000ms
  • Qwen3-TTS 1.7B: about 2500ms / 5300ms

The conclusion was that once you go beyond 200M parameters, it's hard to use as real-time TTS on a Mac. The sweet spot between speed and quality was Kokoro 82M.

The most unexpected finding was that quantization made things slower.

  • fp16 default: fastest at 373ms
  • INT8: 687ms, about 1.8x slower
  • q8f16: 655ms, about 1.75x slower
  • CoreML Neural Engine: architecture unsupported error
  • 4 threads: optimal at about 730ms
  • 1 thread: 1723ms
  • 8 threads: 754ms, which actually introduced overhead

The author concluded that Apple Silicon is optimized for fp16 operations, so quantization like INT8 can reduce memory usage but end up slower due to the cost of type conversion.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.