AI Briefing
KO

RTX 5070 Ti 44 t/s

·2026.04.24 20:02

Key point

On an RTX 5070 Ti 16GB and 32GB RAM, Qwen3.6-35B-A3B Q8_0 achieved 44 t/s.

Details

Running unsloth/Qwen3.6-35B-A3B-GGUF Q8_0 on an RTX 5070 Ti 16GB and 32GB DDR5 RAM setup achieved 44 t/s.

  • Model size: 36.9 GB
  • Context: 128K
  • LM Studio settings:
    • GPU Offload: 40
    • Offload MoE Experts to CPU: 26
    • Try mmap: on
    • K cache: Q8_0
    • V cache: Q8_0

The poster mentioned that llama.cpp would be better.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.