Open-Source GPU Observability Tool l9gpu Released
·2026.05.21 10:49
Key point
l9gpu, an open-source GPU observability tool that can track GPU usage by workload, has been released.
Details
l9gpu is a node-level agent that goes beyond hardware-level metrics, allowing you to see which experiment, team, or job is occupying the GPU.
Key Features:
- Kubernetes Integration: Provides GPU metrics correlated with Pod, Namespace, and Deployment
- Slurm Support: Maps metrics by Job ID, user, and partition
- LLM Inference Optimization: Native metrics support for vLLM, SGLang, and TGI
- Multi-Vendor Hardware Support: Supports NVIDIA, AMD MI300X, and Intel Gaudi
- Standardized Data Output: Exports metrics via OTLP, and includes 17 Prometheus alert rules and a Grafana dashboard
It was developed based on Meta's gcm project, with expanded Kubernetes attribute mapping and multi-vendor support. It is distributed under the MIT license.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.