AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#inference
The latest AI and developer news about #inference, with the original source and a short summary.
Feed
Trending
Tags
Settings
Mica v0.1 4B Released: Open-Source Decision Model for Agent Loops Trained for Under $30
Reddit
·
2026.09.26 07:00
Pick
Open-Weight CLM Released as Jev Alternative, Agent Decision-Making 4x Faster
Reddit
·
2026.09.24 15:00
NVIDIA Releases AIPerf, an LLM Inference Benchmarking Tool
PyTorchKR
·
2026.09.22 18:00
Qwen3.5-based Small Decision Model 'Kev' Released with TypeSafe API Compatibility and Performance Metrics
Hacker News
·
2026.09.21 16:00
tool-prune Released, Tool Schema Selection Speed at 0.4ms
Reddit
·
2026.09.19 01:00
John Hennessy et al. Publish Paper on Intelligence per Watt (IPW) Efficiency of Local LLMs
Hacker News
·
2026.09.14 18:00
AI Can Learn When to Stop Working
Reddit
·
2
·
2026.09.10 16:00
HP Opens Orders for AI Workstation 'ZGX Fury' with GB300 Superchip
Hacker News
·
2026.09.10 04:00
Hugging Face Releases WebGPU Kernels
HuggingFace Blog
·
2026.09.01 09:00
Artificial Analysis Releases Small LLM Inference Benchmark for Mobile Devices
Hacker News
·
2026.08.28 04:00
Sopro V2, Lightweight TTS Released
Reddit
·
2026.08.27 23:00
New Feature Update for ik_llama.cpp
Reddit
·
2026.08.27 16:00
Applied Compute Releases AC2, an AI Model Training Platform for In-House Research
TLDR AI
·
2026.08.26 09:00
kimodo.cpp Released: C++ Port Without PyTorch
PyTorchKR
·
2026.08.25 15:00
Pick
FreeToken Serves 284B MoE Model on RTX 5090
PyTorchKR
·
1
·
2026.08.25 12:00
Up to 30x Work Throughput per Watt: NVIDIA Vera Rubin NVL72 Sets New Record for AI Agent Efficiency
NVIDIA Blog
·
2026.08.25 00:00
LLM Output Compression Cuts Costs by 3x
Reddit
·
2026.08.22 01:00
How to Make Text-to-Speech Models Respond in Under 50ms
Hacker News
·
2026.08.22 00:00
Stealth Model
Hacker News
·
1
·
2026.08.21 08:00
Qwen3.8-27B 2-bit Quantization Released
Reddit
·
2026.08.21 02:00
Syzygy Releases Mach-1 35B
Reddit
·
2026.08.20 16:00
Pick
Cerebras Unveils CS-4
PyTorchKR
·
3
·
2026.08.20 15:00
Breaking Through DeepSeek-V4-Pro Serving Limits
TLDR AI
·
2026.08.20 09:00
DFlash 2: Maintaining Parallel Drafting
Hacker News
·
2026.08.20 05:00
Previous
1
2
3
4
5
Next
Previous
1
2
3
4
5
6
7
Next
#inference | AI Briefing