AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#llm-inference
The latest AI and developer news about #llm-inference, with the original source and a short summary.
Feed
Trending
Tags
Settings
#llm-inference - page 2 | AI Briefing
Comparison of 9 AI Agent Harnesses Based on the Same Model Reveals Up to 17x Cost Difference Per Task
Hacker News
·
2026.09.03 01:00
WebLLM Releases Browser-Based LLM Inference Engine Powered by WebGPU
Hacker News
·
2026.09.02 23:00
NVIDIA Releases Model Lightweighting Tool
PyTorchKR
·
1
·
2026.09.02 09:00
The Efficient Frontier of LLM Inference (6-minute read)
TLDR AI
·
2
·
2026.09.02 09:00
Sliding Window Attention Outperforms Linear Attention
Reddit
·
2026.08.31 23:00
vLLM v0.28.0 Released with Major Optimizations for Kimi-K3 and DeepSeek V4
TLDR AI
·
2026.08.31 09:00
vLLM v0.28.0 Release: Kimi-K3 and DeepSeek V4 Optimizations, Significant Inference Performance Improvements
Hacker News
·
1
·
2026.08.30 03:00
vLLM Fixes Qwen3.8 Non-Determinism Bug
Reddit
·
3
·
2026.08.29 19:00
Qwen3.8-27B SOTA GGUF Released
Reddit
·
1
·
2026.08.29 06:00
Mismatch between GGUF filenames and actual bit counts
Reddit
·
2026.08.29 05:00
Tontaube Releases Open TTS Model
Reddit
·
2026.08.28 18:00
The Pitfalls of Ollama KV Cache Configuration
Reddit
·
2026.08.28 06:00
Samsung Electronics Unveils LPDDR5X-PIM at Hot Chips 2026, Boosting Llama 3.1 Inference Speed by 3x
Hacker News
·
1
·
2026.08.27 04:00
Ollama VRAM Shortage Caused by Default Settings
Reddit
·
2026.08.26 17:00
Qwen3.8-27B NVFP4 Quantization Released
Reddit
·
2026.08.26 10:00
LayerStoRm: Consumer GPU LLM Serving
Reddit
·
2026.08.25 12:00
Hot Chips 2026: High Bandwidth Flash (HBF) Application
Hacker News
·
1
·
2026.08.24 23:00
llama-cpp-turboquant Adds ConvRot Quantization Support
Reddit
·
2026.08.24 10:00
Llama.cpp 0.2.0 Release
Reddit
·
2
·
2026.08.22 15:00
PagedAttention: Virtual Memory for KV Cache
TLDR AI
·
2026.08.21 09:00
LFM2.5 Inference Accelerated 3.2x with DSpark
HuggingFace Blog
·
2026.08.21 01:00
Ollama 0.32.15 Metadata Cache
Reddit
·
2026.08.21 00:00
AirLLM Supports 2.8T Model with 4GB VRAM
Reddit
·
2026.08.20 19:00
Parallel Prefill with ANE+GPU Implemented on Apple Silicon
Reddit
·
2026.08.20 06:00
Previous
1
2
3
4
5
Next
Previous
1
2
3
4
5
6
7
8
9
Next