4070S 12GB Local LLM Speed
·2026.05.01 00:29
Key point
Local inference speeds for Qwen3.6 and Gemma 4 on an RTX 4070S 12GB were shared.
Details
Local model speeds were measured on an RTX 4070S 12GB (+10% OC), Ryzen 9800X3D, DDR5 64GB setup.
Display output was offloaded to the iGPU to save RTX VRAM, using CUDA 13.1 and llama.cpp settings. The core values were n-gpu-layers=999, batch-size=4096, ctx-size=65536, cache-ram=2048, kv-unified=true, flash-attn=true, and Qwen3.6-35B-A3B was run with ctx-size=131072, n-cpu-moe=35.
Results by model are as follows.
Qwen3.6-35B-A3B-GGUF Q6_K_XL: tgs 40, pps 2100Qwen3.6-27B-IQ3_XXS: tgs 16, pps 1000Gemma 4 26B-A4B-it-UD-Q8: tgs 26, pps 2150Gemma-4-31B-it-IQ3_XXS: tgs 13~16, pps 650
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.