Running Qwen3.6-35B on a 6GB VRAM Laptop
·2026.05.04 07:16
Key point
Ran Qwen3.6-35B-A3B at 23 t/s on a 5-year-old laptop with 6GB VRAM.
Details
Hardware
- Asus ROG Zephyrus G14 2020
- Ryzen 7 8c/16t locked at 2,900MHz with boost disabled
- 24GB DDR4-3200
- RTX 2060 Max-Q 6GB
Setup
- Ran
Qwen3.6-35B-A3B-APEX-I-Compact.ggufwithllama-server - 64k context,
--fit off,-fa on,--threads 8,--threads-batch 12 --cpu-range 0-7,--cpu-range-batch 0-11,--cache-type-k/v q8_0,--ubatch-size 1024,--batch-size 2048- Tuned memory and speed with
--cache-ram 4096(4GB) and--spec-type ngram-mod
Results
- Measured throughput was about 23 t/s, maintaining over 10 t/s even while unplugged from power
- Used
pi agentalongside it as a real-world use case
Additional shares
- Also presented a 128k long-context setup using
lm-server-tq, based on Tom's fork
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.