AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#nvfp4
The latest AI and developer news about #nvfp4, with the original source and a short summary.
Feed
Trending
Tags
Settings
Optimizing Gemma 4 Serving with vLLM on Amazon EKS [Part 2: Throughput and Latency SLOs on a Single GPU]
AWS Tech Blog (Korea)
·
2026.09.10 09:00
NInfer Fork Released: Supports 555k Context on RTX 5090 with NVFP4 KV Cache and YARN
Reddit
·
2026.09.06 07:00
NVIDIA Releases NVFP4 Quantized Model of DeepSeek-V4-Pro-0813
TLDR AI
·
2026.08.31 09:00
Qwen3.8-27B NVFP4 Quantization Released
Reddit
·
2026.08.26 10:00
NVFP4 Distillation Preserves Internal Geometry
Reddit
·
2026.08.10 05:00
Laguna-S 2.1 checkpoint released
Reddit
·
2026.08.03 05:00
Laguna S 2.1 Weights Updated
Reddit
·
2026.08.01 22:00
Unsloth Releases NVFP4 Quantization for Qwen3.6
Reddit
·
1
·
2026.07.13 19:00
Optimal Deployment Guide for GLM-5.2 on Blackwell
Reddit
·
2026.07.08 04:00
Pick
NVIDIA Releases Qwen3.6 NVFP4
Reddit
·
2026.05.31 02:00
llama.cpp Adds NVFP4/MTP Support
Reddit
·
2026.05.24 03:00
GitHub repository for real-time long video generation
TLDR AI
·
2026.05.20 09:00
NVIDIA Releases Kimi NVFP4
Reddit
·
1
·
2026.05.14 21:00
Mac Mini M4 Pro 70 t/s
Reddit
·
2026.05.03 14:00
Gemma 4 26B NVFP4 Released
Reddit
·
2026.05.01 12:00
llama.cpp NVFP4 Comparison
Reddit
·
2026.04.29 21:00
llama.cpp Adds Blackwell NVFP4 Support
Reddit
·
2026.04.29 17:00
llama.cpp merges SM120 NVFP4 MMQ
Reddit
·
2026.04.29 09:00
GLM 5.1 running locally at 40tps
Reddit
·
2026.04.26 01:00
llama.cpp and ik_llama.cpp FP4 Support
Reddit
·
1
·
2026.04.26 00:00
Qwen3.6-27B 80 tps
Reddit
·
2026.04.25 19:00
#nvfp4 | AI Briefing