Lemonade v10.8 Released: Turning Local Models into MCP Tools
·2026.06.18 04:42
Key point
Lemonade v10.8 introduces automatic memory management, cloud offload, and MCP gateway functionality.
Details
The Lemonade v10.8 update adds intelligent memory and context management features. It now dynamically manages VRAM to automatically unload idle models and adjust KV-cache size to free up GPU memory, and model pinning lets you keep required models resident in memory.
Key updates include the following:
- Cloud Offload: Supports a provider-agnostic backend that lets you use OpenAI-compatible providers (Fireworks, OpenRouter, Together, etc.) alongside local models.
- MCP Gateway: Exposes local models as MCP (Model Context Protocol) tools, allowing MCP-supporting hosts to call local models as tools instead of cloud APIs.
- LMX-Omni Expansion: Provides controls such as size and step count for image generation, and allows importing custom Omni models from Hugging Face.
- Platform Expansion: Strengthened support for various hardware and OS, including NVIDIA Blackwell (GB10), AMD ROCm (Windows/Linux), and Debian 13 builds.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.