AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#gpu-inference
The latest AI and developer news about #gpu-inference, with the original source and a short summary.
Feed
Trending
Tags
Settings
Pick
Netflix's In-House LLM Serving Infrastructure
Netflix Tech
·
1
·
2026.07.18 06:00
llama.cpp Adds MTP Support
Reddit
·
2026.05.19 04:00
Qwen3.6 27B on Dual RTX 5060 Ti, 60 tok/s
Reddit
·
2026.04.29 17:00
22 tok/s on RTX 5060 Ti 16GB
Reddit
·
2026.04.24 09:00
85TPS on a single 3090
Reddit
·
2026.04.23 23:00
Zero-Copy GPU Inference with WebAssembly on Apple Silicon
Hacker News
·
2026.04.19 07:00
MIG cuts p95 in half
Reddit
·
2026.04.18 04:00
40 tok/s on a 3080
Reddit
·
2026.04.17 05:00
2x Speed with WSL
Reddit
·
2026.04.16 11:00
122B 198t/s
Reddit
·
2026.04.10 09:00
Beyond Two-Tower: Redesigning the Serving Stack for Next-Generation Lightweight Ad Ranking Models
Pinterest Engineering
·
2026.02.03 02:00
#gpu-inference | AI Briefing