AI Briefing
KO

Qwen3.6-27B 80 tps

·2026.04.25 19:21

Key point

The author reported running Qwen3.6-27B on a single RTX 5090 with vLLM 0.19.1rc1, achieving about 80 tokens/sec at a 218k context.

Details

Served Qwen3.6-27B on a single RTX 5090 using vLLM 0.19.1rc1, achieving about 80 tps.

  • Context window: 218k
  • Model/weights: Qwen3.6-27B Text NVFP4 MTP version
  • Serving stack: latest vLLM 0.19 build
  • Reproducibility note: explained that the same recipe from the previously posted Qwen3.5-27B setup was applied as-is.

Linked both the NVFP4 MTP checkpoint uploaded on HF and the earlier Qwen3.5-27B test post, suggesting the performance can be reproduced the same way.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.