AI Briefing
KO

Spark-X2.5-4B-GGUF: GGUF Distribution of the 4B Lightweight LLM with 1M Context Support

XHToken/Spark-X2.5-4B-GGUF

·2026.09.11 02:36

This is the BF16 GGUF conversion of the Spark-X2.5-4B model, enabling lightweight inference in local environments. It is a small language model designed for general-purpose tasks such as conversation, code generation, translation, reasoning, and agent workflows.

It features a hybrid attention architecture supporting a native context length of up to 1M tokens and handles over 200 languages. A thinking mode toggle allows flexible adjustment between response speed and reasoning depth.

Optimized for compatibility with major local execution environments such as llama.cpp, Ollama, and LM Studio. Released under the Apache 2.0 license, it permits free commercial use and customization, and offers various quantization versions including Q4_K_M.

HuggingFace
HuggingFace model

XHToken/Spark-X2.5-4B-GGUF

The original page has no description.

text-generation

This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.