AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#llm-inference
The latest AI and developer news about #llm-inference, with the original source and a short summary.
Feed
Trending
Tags
Settings
#llm-inference - page 3 | AI Briefing
llama.cpp Releases v0.1.0 with Semantic Versioning
Reddit
·
1
·
2026.08.17 22:00
Qwen 3.8 Achieves 288k Tokens/s on GB300
Reddit
·
2026.08.17 02:00
bitsandbytes Creator Previews New Quantization Method
Reddit
·
2026.08.14 22:00
Unsloth Desktop - Open-source app unifying local AI model execution, training, and agents
GeekNews
·
2026.08.13 10:00
WARP Runs K3 on 64GB MacBook
PyTorchKR
·
1
·
2026.08.12 18:00
llama.cpp LLM Inference Accelerated 11–16x on Apple Silicon macOS VMs
GeekNews
·
2026.08.12 05:00
Ling-3.0-tiny Released
Reddit
·
2026.08.11 02:00
Kimi K3 GGUF Released
Reddit
·
2026.08.07 17:00
Beyond Next-Token Prediction: Performance Characteristics of Diffusion and Autoregressive Language Models
Apple ML
·
2026.08.07 09:00
Arbitrage: Efficient Inference via Speculative Decoding Considering Model Superiority
Apple ML
·
2026.08.07 09:00
DiffusionGemma Achieves 1,500 Tokens/s
PyTorchKR
·
2026.08.06 15:00
Core Technologies for LLM Inference Optimization: Quantization, KV Cache, and Inference Chips
KT Cloud
·
2026.08.06 14:00
Ling-3.0-Flash Released
Reddit
·
1
·
2026.08.05 12:00
Running 70B Model Inference on a Single 4GB GPU with AirLLM
Hacker News
·
2026.08.03 20:00
WASTE Runs Ultra-Large Models with NVMe
Reddit
·
2026.08.03 09:00
Predictive Speculative KV Replication for Bursty LLM Inference
Hacker News
·
2026.08.01 04:00
Everyone is building an LLM router, but we killed ours
Hacker News
·
1
·
2026.08.01 03:00
Running Kimi K3 at 0.50 tok/s with 29GB RAM
Hacker News
·
2026.07.31 23:00
Fermion Research Unveils Neutrino-1 Model Family
PyTorchKR
·
2026.07.30 09:00
EschaLabs Releases 2-bit Qwen3.6-35B-A3B Quantized Model
Reddit
·
2026.07.30 06:00
AI Gateway Adds Unified Fast Mode Support
Vercel Blog
·
2026.07.30 03:00
Launch HN: Tokenless (YC S26) – Automatic Model Switching to Cut Costs
Hacker News
·
2026.07.30 00:00
Kimi K3 and Kimi K3 Fast Now Support ZDR and Feature US-Based Provider AI Gateway
Vercel Blog
·
2026.07.27 09:00
How to Build an Ultra-Fast API for GLM-5.2
TLDR AI
·
2026.07.27 09:00
Previous
1
2
3
4
5
Next
Previous
1
2
3
4
5
6
7
8
9
Next