AI Briefing
KO

NVFP4 Distillation Preserves Internal Geometry

·2026.08.10 05:22

Key point

CKA-QAD improves inference and coding performance by preserving internal representations during NVFP4 quantization-aware distillation.

Details

Analysis suggests that in NVFP4 Quantization-Aware Distillation (QAD) for low-bit inference, matching only the output distribution can significantly degrade the model's internal representations.

Researchers measured layer-wise representation similarity between BF16 teacher models and quantized student models using CKA, confirming that QAD using only KL loss causes the geometry of internal representations to collapse. Drift was particularly severe in post-reinforcement-learning models, which was linked to performance bottlenecks in inference and coding tasks.

To address this, the authors proposed the CKA-QAD regularization technique, which aligns layer-wise Gram matrices. Experiments with Nemotron 3 Nano and Qwen3-4B-Thinking-2507 demonstrated that aligning internal representations improves downstream inference and coding accuracy while keeping training overhead modest.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.