AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#llm-inference
The latest AI and developer news about #llm-inference, with the original source and a short summary.
Feed
Trending
Tags
Settings
#llm-inference - page 5 | AI Briefing
MBD-LM: Optimizing Parallel Generation for Diffusion-Based Language Models
Reddit
·
2026.07.04 22:00
New 'Scatter' sampler added to llama.cpp
Reddit
·
2026.07.04 06:00
ReFreeKV: Adaptive KV Cache Compression Technique
Reddit
·
2026.07.03 23:00
Program-as-Weights: A Programming Paradigm for Fuzzy Functions
Hacker News
·
2026.07.03 21:00
Learning Unmasking Policies for Diffusion Language Models
Apple ML
·
2026.07.02 09:00
DeepSeek Open-Sources New Framework 'DSpark' That Boosts LLM Inference Speed by Up to 85%
TLDR AI
·
2026.06.30 09:00
llama.cpp Begins Supporting DeepSeek V4
Reddit
·
2026.06.30 03:00
ggml CUDA Performance and Synchronization Optimization
Reddit
·
2026.06.27 13:00
JetSpec Boosts LLM Inference Speed by Up to 9.64x
Reddit
·
2026.06.26 06:00
OpenAI and Broadcom Unveil Chip Optimized for LLM Inference
OpenAI Blog
·
2026.06.24 15:00
Data-Driven Analysis of the Local AI Ecosystem Based on r/LocalLLaMA
Reddit
·
2026.06.20 17:00
Lemonade v10.8 Released: Turning Local Models into MCP Tools
Reddit
·
1
·
2026.06.18 04:00
DFlash and Spec V2 Decoding Technology
TLDR AI
·
2026.06.16 09:00
NVIDIA Blackwell Takes First Place in the First Agentic AI Infrastructure Benchmark
NVIDIA Blog
·
1
·
2026.06.13 06:00
Could You Buy Your KV Cache
Hacker News
·
1
·
2026.06.13 05:00
MiniMax Unveils MSA for Ultra-Long-Context Processing
Reddit
·
2026.06.12 23:00
EAGLE3 integration into llama.cpp completed
Reddit
·
2026.06.12 16:00
InfiniteKV: Open-Source KV Cache for Infinite Context
Reddit
·
2026.06.12 15:00
Pick
DiffusionGemma: 4x Faster Text Generation
Google AI Blog
·
1
·
2026.06.11 01:00
FlashMemory DeepSeek-V4 Retriever (GitHub Repo)
TLDR AI
·
2026.06.10 09:00
OSCAR: 2-bit KV Cache Quantization Technique Released
Reddit
·
2026.06.10 04:00
Nvidia Proposes High-Performance CPU System for Windows PCs
GeekNews
·
2026.06.07 09:00
Domino: 5.8x Faster Inference
Reddit
·
2026.06.06 21:00
llama.cpp Dynamic KV Cache Quantization Feature
Reddit
·
2026.06.05 03:00
Previous
3
4
5
6
7
Next
Previous
1
2
3
4
5
6
7
8
9
Next