Spark-X2.5-4B-GGUF: GGUF Distribution of the 4B Lightweight LLM with 1M Context Support
XHToken/Spark-X2.5-4B-GGUF
About the project
This is the BF16 GGUF conversion of the Spark-X2.5-4B model, enabling lightweight inference in local environments. It is a small language model designed for general-purpose tasks such as conversation, code generation, translation, reasoning, and agent workflows.
It features a hybrid attention architecture supporting a native context length of up to 1M tokens and handles over 200 languages. A thinking mode toggle allows flexible adjustment between response speed and reasoning depth.
Optimized for compatibility with major local execution environments such as llama.cpp, Ollama, and LM Studio. Released under the Apache 2.0 license, it permits free commercial use and customization, and offers various quantization versions including Q4_K_M.
XHToken/Spark-X2.5-4B-GGUF
The original page has no description.
text-generation
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.