AI Briefing
KOSign in

Bonsai Ternary Model Enables 27B-Class Reasoning on Consumer Hardware

·2026.10.11 21:58

Key point

The model retains 98.2% of FP16 intelligence while reducing size to ~5.9 GB for llama.cpp.

Details

Bonsai introduces a ternary transformer architecture that enables full 27B-class reasoning on standard laptops and single GPUs, addressing the needs of GPU-poor users. The model achieves ~9.3x size reduction compared to FP16, shrinking from ~54 GB to ~5.9 GB, while retaining 98.2% of FP16 intelligence.

Performance and Benchmarks

The model demonstrates strong retention of reasoning capabilities in the sub-4-bit regime, where conventional low-bit representations typically collapse:

  • Average Score: 84.78 across 14 thinking-mode benchmarks, significantly outperforming the conventional IQ2_XXS build (72.59) at a similar footprint.
  • Math: 96.57, within half a point of full precision.
  • Coding: 89.42, level with the baseline.
  • Agentic Tool Calling: 74.92.
  • Speed: Approximately 47 tokens per second on an Apple M5 Max laptop.

Technical Implementation

Bonsai uses end-to-end ternary language weights across embeddings, attention projections, MLP projections, and the LM head, achieving a true 1.72 bits per weight without high-precision escape hatches. The vision tower is shipped separately as a Q8_0 mmproj pack.

The model is built on the Qwen3.8-27B hybrid-attention backbone (~75% linear attention), supporting a 262K-token context on-device. It offers two GGUF packings for llama.cpp (CUDA, Metal):

  • PTQ1_0: Packs trits densely at 1.75 bits/weight (5.95 GB).
  • PQ2_0: Stores each trit in a 2-bit slot at 2.13 bits/weight (7.21 GB).

Packed weights are consumed directly by custom kernels and are never expanded back to FP16.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.