AI Briefing
KO

Dual GX10 Benchmark

·2026.04.21 19:36

Key point

This is a benchmark result running MiniMax M2.7 4bit on dual Asus GX10 at about 100W.

Details

Ran MiniMax-M2.7-AWQ-4bit without issues on a dual Asus GX10 (dgx spark) setup.

  • Power consumption per GX10 during inference is around about 100W
  • Works without errors in open code and hermes agent
  • Time to first token feels long, but the low-noise, low-heat setup is an advantage

Performance based on llama benchy --depth 0 4096 8192 16384 32768 --latency-mode generation is as follows.

  • pp2048: 3452.05 t/s, ttfr 626.82 ms
  • tg32: 38.84 t/s
  • pp2048 @ d4096: 2848.85 t/s, ttfr 2022.61 ms
  • pp2048 @ d8192: 2579.85 t/s, ttfr 3523.69 ms
  • pp2048 @ d16384: 2411.34 t/s, ttfr 6791.62 ms
  • pp2048 @ d32768: 1988.05 t/s, ttfr 15512.61 ms

As context length increases, prompt processing speed drops, but generation speed stayed in the range of about 37-39 t/s.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.