AI Briefing
Sign in

vllm-xtu-moe enables 748B MoE inference on 2×A100-40GB via CPU-GPU fused execution

·2026.09.29 14:18

Key point

The open-source vllm-xtu-moe project achieves up to ~20× faster prefill than CPU-only paths for medium contexts by streaming expert weights from CPU to GPU.

Details

The vllm-xtu-moe project enables the execution of a 748B MoE model on 2×A100-40GB GPUs by keeping routed expert weights out of VRAM. The system slices expert weights along the CPU's physical topology, ensuring one copy in memory with node-local reads via NUMA binding and first-touch, allowing the engine to operate near the machine's aggregate DRAM bandwidth.

Performance and Architecture

The architecture supports AMD EPYC nps=4 and is built for SM 8.x compute capabilities, including the 30-series family (SM86). Long prompts stream weights to the GPU using double buffering (ping/pong) overlapped with attention, achieving up to ~20× faster prefill than the CPU path at medium context and 2–3× at long context. Short prompts are prefilled directly on the CPU.

VRAM allocation follows a fixed priority: KV pool → GPU-prefill staging → draft weights → activation workspace. If VRAM is insufficient, GPU prefill is dropped first, falling back to pure CPU prefill. Speculative decoding is enabled when space permits; MiMo-V2.6 MTP k=1 measures 83% first-position acceptance and +19% decode speedup (with a DSpark anchor at +20%).

Benchmarks

Measurements were conducted on 2×A100-40GB with TP=2 using vllm bench serve:

  • GLM-5.3-Flash: 115.2 tok/s prefill (140 prompt, C=1), 22.0 tok/s decode.
  • DeepSeek-V4.1-Flash: 903.9 tok/s prefill (16,384 prompt, C=1), 19.3 tok/s decode.
  • MiMo-V2.6-Flash-RL: 1,455.3 tok/s prefill (16,384 prompt, C=2), 39.6 tok/s decode.

The system saturates CPU prefill at ~33k expert-tokens/s and supports 704K max context with bf16 KV. It remains a patch series and plugin for mainline vLLM under the Apache-2.0 license.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.