AI Briefing
KO

64 tok/s peak on 8x MI50

·2026.04.17 02:39

Key point

A benchmark running MiniMax M2.7 AWQ on 8x MI50 32GB hit a peak of 64 tok/s.

Details

This is the result of running a benchmark after serving MiniMax-M2.7-AWQ-4bit on an 8x MI50 32GB setup using the vllm-gfx906-mobydick fork.

  • Quantized model used: cyankiwi/MiniMax-M2.7-AWQ-4bit
  • Engine used: vllm-gfx906-mobydick fork
  • Runtime environment: ROCm-based Docker, tensor parallel size 8, max context length 196,608

The benchmark settings were random input of 10,000 tokens, output of 1,000 tokens, 4 prompts, and a request rate of 10,000 RPS.

The results were recorded as 4 successful / 0 failed, a total of 40,000 input tokens and 400 generated tokens, with a benchmark duration of 125.90 seconds.

The key points are a real-world working example of the ROCm vLLM fork running on gfx906-series MI50 hardware, and the 64 tok/s peak performance achieved with this combination.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.