AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#llm-inference
The latest AI and developer news about #llm-inference, with the original source and a short summary.
Feed
Trending
Tags
Settings
#llm-inference - page 6 | AI Briefing
DeepLearning.AI Launches Official vLLM Course
Reddit
·
2026.06.05 02:00
Speculative KV Coding: Losslessly Compressing the KV Cache by Up to ~4x
Hacker News
·
2026.06.05 00:00
OCTOPUS: Optimizing Transformer KV Cache via Optimal Squared-Error Quantization and Octahedral Parameterization
Stability AI Research
·
2026.06.04 01:00
llama.cpp adds reasoning level control feature
Reddit
·
1
·
2026.06.02 22:00
llama.cpp adds Step-3.7-Flash support
Reddit
·
2026.06.02 18:00
llama.cpp Proposes VRAM Usage Optimization PR
Reddit
·
2026.06.02 00:00
llama.cpp Adds Support for EXAONE 4.5
Reddit
·
2026.06.01 18:00
llama.cpp Unified Binary and Website
Reddit
·
2026.05.30 01:00
Real-Time LLM Inference on Standard GPUs: 3,000 tokens/s per Request
Hacker News
·
2026.05.29 18:00
AMD MI300X LLM Inference Optimization Technology
Reddit
·
2026.05.29 17:00
llama.cpp Optimizes VRAM Usage for Flash Attention
Reddit
·
2026.05.29 16:00
llama.cpp AMD/ROCm Update
Reddit
·
1
·
2026.05.29 10:00
Zai Optimizes GLM-5.1 Inference Network
Reddit
·
1
·
2026.05.28 22:00
Zai Optimizes GLM-5.1 Inference with ZCube Adoption
Reddit
·
2026.05.28 22:00
GPU Job Scheduling Using an Idle Inference GPU Pool
GeekNews
·
1
·
2026.05.27 08:00
TogetherAI Unveils OSCAR KV Quantization
Reddit
·
2026.05.26 21:00
2-bit KV Cache Quantization Technique 'OSCAR' Released
Reddit
·
2026.05.25 20:00
llama.cpp Optimizes Context Re-processing
Reddit
·
1
·
2026.05.25 15:00
AMD RDNA3-Optimized Inference Engine hipEngine
Reddit
·
2026.05.25 07:00
llama.cpp Asymmetric KV Cache Quantization Issue and Optimization Discussion
Reddit
·
3
·
2026.05.22 22:00
Mix-Quant: Quantize the Prefill, Keep the Decoding at High Precision
Reddit
·
2026.05.22 04:00
Open-Source GPU Observability Tool l9gpu Released
Reddit
·
2026.05.21 10:00
AWS Offers M3 Ultra Mac Studio
Reddit
·
2026.05.20 23:00
llama.cpp Improves Prompt Processing Speed for MTP
Reddit
·
2026.05.18 00:00
Previous
4
5
6
7
8
Next
Previous
1
2
3
4
5
6
7
8
9
Next