290MB Local LLM
·2026.04.16 01:29
Key point
The 290MB 1-bit Bonsai 1.7B runs locally in the browser via WebGPU.
Details
A demo has been released that shrinks 1-bit Bonsai 1.7B down to 290MB and runs it locally inside the browser using WebGPU.
- Model size: 290MB
- Execution method: Local execution in browser
- Acceleration: WebGPU
- Demo link:
bonsai-webgpuon Hugging Face Spaces
In the comments, there were responses wanting to see more of the actual tokens/s performance of such ultra-lightweight 1-bit models based on llama.cpp, and CPU, Metal, and Vulkan support were mentioned as currently available. CUDA support is said to be in progress.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.