Gemma 4 breakthrough
Key point
Ran Gemma 4:26B on a Mac Studio M4 Max using OC + LM Studio, achieving 40 tok/s.
Details
After adjusting several settings to run Gemma 4:26B on a Mac Studio M4 Max (36GB RAM, 512GB SSD), the author succeeded in getting it working with a combination of LM Studio and a separate OpenClaude (OC) instance.
The key was installing Docker Desktop to run OC inside a sandbox container, and there was no noticeable latency or slowdown.
Performance came in at around 40 tokens/s, and while context up to 64k is being tested, raising it to 262k caused errors.
Additional settings applied were as follows.
- K/V cache quantization: Q4_0
- Enabled the option to keep the model in memory
The author stated they plan to keep testing how strong this setup is compared to models like Qwen 3.5 or GLM 4.7 Flash.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.