Power Limits and TG/s on 2x3090
·2026.04.28 13:48
Key point
On 2x3090, a 250W power limit emerged as the efficiency sweet spot for Qwen3.6-27B.
Details
Serving Qwen3.6-27B on 2 RTX 3090 GPUs, the relationship between power limit and TG/s was compared.
- Reported that 250W gave the best balance between performance and power efficiency.
- Under a 1 concurrent request condition, higher TG/s was achieved at 275W.
- The experiment included vLLM server settings and benchmark commands, increasing reproducibility.
- The configuration included tensor parallel size 2, prefix caching, speculative decoding, fp8 KV cache, chunked prefill, and more.
Local LLM inference throughput can vary depending on the power limit, and these figures are worth referencing when finding the efficiency sweet spot in a 2x3090 environment.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.