Edge0-35B-A3B-preview: Edge Inference Optimization for 35B MoE Models
Edge0/Edge0-35B-A3B-preview
About the project
This model is optimized for edge device environments with a 35B MoE architecture. Although it has a total of 35B parameters, it maintains approximately 3B active parameters during inference to reduce computational load. By applying 4-bit quantization and SSD offloading techniques, it delivers the performance of large-scale models even on limited hardware.

Designed to run on Apple Silicon environments based on the MLX framework, it includes a separate file for prompt routing, allowing flexible adjustment of model behavior according to request complexity. It also supports LoRA adapters for lightweight custom tuning.
Specialized for text generation and conversational tasks, it is distributed under the Apache 2.0 license. You can launch a local server to integrate via an OpenAI-compatible API or perform direct inference through MLX LM. It is suitable for developers who need to run high-performance LLMs on desktops or laptops.
Edge0/Edge0-35B-A3B-preview
The original page has no description.
text-generation
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.