AI Briefing
KO

INT3·INT2 KV Cache Released

·2026.04.22 15:54

Key point

A Metal kernel implementation combining INT3 model compression with INT2 KV cache has been released.

Details

An implementation combining an INT3 quantized model with INT2 KV cache has been released.

  • The compression loss is reported to be around +0.14 nats.
  • Custom fused Metal kernels for Mac M-series were used to jointly optimize the model and the KV cache.
  • Qwen 7B is currently available as a preview.
  • Installation is available via brew install reinforceai/spiral/spiral, and it can be run with spiral-chat.
  • The author announced plans to add GPU support later using Triton kernels, along with more efficient packing and additional model releases.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.