AI Briefing
KO

WASTE Runs Ultra-Large Models with NVMe

·2026.08.03 09:16

Key point

WASTE streams activated expert weights from NVMe, enabling Kimi K3 to run larger than RAM.

Details

WASTE is an embedded C-based inference engine that operates without any third-party runtime dependencies.

It keeps the model's trunk in memory while streaming only the selected expert weights directly from NVMe in MoE models. The remaining RAM is used as a size-limited expert cache.

Through this structure, it aims to run the 2.78 trillion parameter Kimi K3 model, which exceeds available RAM capacity. However, the post does not provide additional operational information such as performance, supported hardware, or licensing.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.