AI Briefing
KO

APEX MoE Quantized Models Expand to Over 30 Types

·2026.05.05 01:43

Key point

APEX, a mixed-precision quantization strategy for MoE models, has expanded to over 30 model types and launched a new ultra-compressed tier.

Details

This is an update on APEX, a mixed-precision quantization strategy for MoE (Mixture-of-Experts) models. Following the existing Qwen 3.5 model, over 30 major MoE models are now available quantized using the APEX method.

Key Features and Feedback:

  • Long-context retention: APEX's I-Balanced and I-Compact methods maintain higher precision for shared experts and edge layers. As a result, they show high consistency even at 32k tokens or longer contexts, without the performance degradation typically seen with standard Q4_K quantization.
  • Strong coding performance: For the Qwen 3.6 35B-A3B model, the I-Compact and I-Mini tiers show performance close to the FP16 model in real-world coding tasks.

Newly Added Key Models:

  • Qwen family: Qwen 3.5 (122B, 397B, etc.), Qwen 3.6 (35B-A3B), Qwen3-Coder, and more
  • Frontier-class MoE: MiniMax-M2.5/M2.7 (up to 228B), Mistral-Small 4 (119B), NVIDIA Nemotron-3-Super (120B), GLM-4.7 Flash, Step-3.5 Flash, Nemotron-3-Nano (with multimodal support), and more

With this update, the I-Mini/I-Compact tiers, which can run on consumer GPUs, are expected to see even greater use.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.