With 96GB, you can go up to 122B
Key point
For 96GB VRAM setups, Qwen 3.5 122B and 27B have emerged as the practical choices.
Details
In a 96GB VRAM environment, Qwen 3.5 122B was cited as a realistic upper limit. There was an experience report that quantizing with AWQ / GPTQ / NVFP4 enables 110+ tps on vLLM even at 200K+ context.
On the other hand, there was an opinion that Qwen 27B comes close to the 122B MoE on some tasks, but feels slower in practice. It was also pointed out that 96GB is a somewhat awkward range — not enough to run large models very comfortably, yet more than enough for mid-sized models.
Real-world usage examples included:
- Running Qwen3.5-122B-A10B on vLLM with tensor parallel 4
- Boosting efficiency with settings like
max-num-seqs 2,max-num-batched-tokens 4096, andattention-backend flashinfer - An assessment that AWQ is better than
Intel/...AutoRoundin terms of TP=4 support and quality
Another response summarized that 96GB allows up to 120B+ with a high KV cache. On the image generation side, there was also a mention of HunyuanImage-3 running smoothly on 96GB-class hardware.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.