llama.cpp b9180 Adds MTP Support
Key point
The llama.cpp b9180 release adds MTP support and partial rollback.
Details
The llama.cpp b9180 release is out, including Multi-Token Prediction (MTP) support.
Centered on spec: support MTP, it includes cleanup of draft-mtp naming, batch size fixes, and conversion/documentation/format improvements.
GDN partial rollback has been added to the speculative decoding path, allowing rollback up to draft_max even when some draft tokens are rejected.
Server and memory handling were also improved.
- RS-based MTP is disabled when used together with other spec types
- Checkpoint logic adjustments and early-exit fixes
- Added tests for dirty ctx loading
Release artifacts are provided for macOS, Linux, Android, Windows, and openEuler.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.