Super-accelerating with 2x 3090
Key point
Ran Qwen3.6 35B 4bit for multi-user use via vLLM Docker on 2x RTX 3090.
Details
On a 2x RTX 3090 server, loaded cyankiwi/Qwen3.6-35B-A3B-AWQ-4bit using the vLLM OpenAI Docker image to improve multi-user response performance.
The configuration is as follows.
tensor-parallel-size 2max-model-len 65536gpu-memory-utilization 0.85enable-prefix-cachingreasoning-parser qwen3enable-auto-tool-choicetool-call-parser qwen3_codermax-num-seqs 32speculative-config:qwen3_next_mtp, speculative tokens 2
In the benchmark, pp2048 @ d2000 recorded 5463.38 t/s ± 111.87, and tg32 @ d2000 recorded 103.13 t/s ± 22.06.
Another test, pp2048 @ d32768, also produced 5178.25 t/s, showing that a 35B-class quantized model can be operated with high throughput on a 2x 3090 setup.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.