AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#quantization
The latest AI and developer news about #quantization, with the original source and a short summary.
Feed
Trending
Tags
Settings
#quantization - page 4 | AI Briefing
llama.cpp Dynamic KV Cache Quantization Feature
Reddit
·
2026.06.05 03:00
DeepLearning.AI Launches Official vLLM Course
Reddit
·
2026.06.05 02:00
KVarN: A vLLM-Native KV-Cache Quantization Backend Developed by Huawei
Hacker News
·
2026.06.05 00:00
OCTOPUS: Optimizing Transformer KV Cache via Optimal Squared-Error Quantization and Octahedral Parameterization
Stability AI Research
·
2026.06.04 01:00
Bonsai Image 4B: Ultra-lightweight Diffusion Transformer Released
Reddit
·
2026.06.02 23:00
Holo3.1: Local/Mobile Agent Released
HuggingFace Blog
·
2026.06.02 23:00
llama.cpp Fixes Multi-GPU KV Cache Bug
Reddit
·
2026.06.02 05:00
1-Bit Bonsai Image 4B Image Generation Model for Local Devices
TLDR AI
·
2026.06.01 01:00
GGUF Release with Integrated MTP Head
Reddit
·
1
·
2026.05.31 14:00
Pick
NVIDIA Releases Qwen3.6 NVFP4
Reddit
·
2026.05.31 02:00
vLLM Adds HIP W4A16 Kernel Support for AMD
Reddit
·
2026.05.29 21:00
PrismML Releases Ultra-Lightweight Bonsai Image 4B
Reddit
·
2026.05.27 03:00
Shard Released, Compressing KV Cache by 10x
Reddit
·
2026.05.26 13:00
llama.cpp Improves Inference Speed with CUDA-based FWHT
Reddit
·
2026.05.26 02:00
2-bit KV Cache Quantization Technique 'OSCAR' Released
Reddit
·
2026.05.25 20:00
W8A8 Quantization SDK 'Cider' for MLX Released
Reddit
·
2026.05.25 17:00
1.58-bit LLM Training System for Ascend NPUs
Reddit
·
2026.05.25 00:00
llama.cpp Adds NVFP4/MTP Support
Reddit
·
2026.05.24 03:00
llama.cpp Asymmetric KV Cache Quantization Issue and Optimization Discussion
Reddit
·
3
·
2026.05.22 22:00
Cohere Unveils Apache 2.0 Open Model 'Command A+'
Reddit
·
2026.05.22 13:00
Mix-Quant: Quantize the Prefill, Keep the Decoding at High Precision
Reddit
·
2026.05.22 04:00
Tencent Unveils Multilingual Translation Model Hy-MT2
Reddit
·
2026.05.21 21:00
Cohere Unveils Command A+
Reddit
·
2026.05.21 06:00
Qwen 3.6 35B GGUF Released
Reddit
·
2026.05.21 00:00
Previous
2
3
4
5
6
Next
Previous
1
2
3
4
5
6
7
8
9
10
Next