SuperGemma4 - An Uncensored/Speed-Improved/Quantized Model of Google Gemma 4 26B
Key point
A fast, uncensored local model optimized from Gemma 4 26B into MLX 4-bit.
Details
An Apple Silicon MLX-optimized 4-bit quantized text-only model based on Gemma 4 26B IT. Its size is about 13GB.
It aims to be smarter than the original, run faster on the same machine, and deliver stable output for code generation, tool use, and Korean-language prompts.
Key figures are also presented.
- Quickbench score of 95.8, up from the original 91.4
- Generation speed of 46.2 tok/s, about 8.7% faster
- Code generation score of 98.6, +6.3 over the original
- Korean prompt score of 95.0, +4.3 over the original
It particularly emphasizes that output remains stable while retaining its uncensored nature. It states there are no answers blocked by content filters, and that it is more stable in code generation and agent-style prompts as well.
It is also introduced as ready to deploy for local agent workloads.
- Browser automation
- Tool calling
- Planning
- Text processing for local pipelines
An example of running it is as follows.
mlx_lm.server --model Jiunsong/supergemma4-26b-uncensored-mlx-4bit-v2 --port 8080
It automatically supports OpenAI-compatible serving, and no separate template configuration is said to be needed. In fact, it warns that putting a path in --chat-template may corrupt responses.
The format is MLX 4-bit, BF16/U32 tensors, and Safetensors.
In the comments, there was a point raised that the license differs from the original Gemma 4, along with a mention that it is not Apache 2.0.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.