AI Briefing
KO

122B MoE Local Inference with 8GB VRAM

·2026.06.04 03:44

Key point

InstinctRazor technology has been unveiled, running a 122B MoE model even in an 8GB VRAM environment by placing experts on the CPU.

Details

InstinctRazor-Qwen3.5-122B-A10B is a 122B MoE (Mixture of Experts) model and runtime configuration that dramatically lowers GPU VRAM requirements by keeping Experts on the CPU.

Through this approach, users can run local inference on large models with only 8GB-level consumer GPU VRAM. The total compressed model size is approximately 50GB.

Key Benchmark Performance (vs. Gemma-4-A4B):

  • Wins: MMLU-Pro, GPQA-Diamond, MMMLU, HLE, LiveCodeBench
  • Losses: MATH-500, AIME

The project focuses on optimizing the tradeoff between memory usage and execution time.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.