0.27.11mo ago
vLLM v0.27.1
Key point
vLLM v0.27.1 is a patch release that supports quantized DSpark Markov heads.
Details
- v0.27.1 is a patch release based on v0.27.0.
Key Features
- Added support for quantized DSpark Markov heads. (#50424)
vllm-project/vllm
vLLM is a fast and easy-to-use library for large language model (LLM) inference and serving. It is primarily used by ML engineers deploying models as services, platform/infrastructure engineers, and researchers, and is utilized in production environments thanks to its OpenAI-compatible API server and support for various models and hardware. Releases are important because they directly impact performance, supported models, hardware compatibility, and API behavior.
vLLM v0.27.1 is a patch release that supports quantized DSpark Markov heads.