AI Briefing
KO

UkisAI Releases Swift Series: Efficient Qwen-Based LLMs with Reduced Thinking Tokens

·2026.09.25 01:32

Key point

The new Swift 1.5 27B model reduces thinking tokens by 58.5% while improving accuracy, with Flash Next offering a 1.8x speed increase.

Details

UkisAI has released the Swift series, a family of efficient reasoning LLMs based on Qwen, designed to mitigate pathological overthinking patterns. The models utilize GSQ-RCO quantization and are trained by penalizing excessive thinking tokens while restoring accuracy via RL (GSPO) and OPD.

Model Performance and Metrics

The release includes three primary variants, each benchmarked across General (GPQA, AIME26), Coding (LiveCodeBench), Vision (ERQA), and Agentic (Terminal Bench 2.1) domains, averaged over five seeds:

  • Swift 1.5 27B: Achieves a 58.5% reduction in thinking tokens while scoring 0.35% higher than the base model. It outperforms the base on Terminal Bench 2.1 by avoiding "overthinking error" loops, though this results in higher average token usage for successful task completion. The source notes the Terminal Bench score is misleadingly low due to this behavior, with an apples-to-apples token reduction of -38.7%.
  • Swift Flash Next: Delivers a 63.4% reduction in thinking tokens and a 1.8x speedup, with a negligible -0.2% accuracy difference compared to the base model at xhigh settings.
  • Swift Bonsai 2: An experimental variant that reduces thinking tokens by 39.8% while scoring 0.19% higher than the base.

Availability and Infrastructure

The models are available on Hugging Face with various quantizations, including GGUF, NVFP4, MLX, and W4A16. UkisAI also provides a Research API and Hugging Face Spaces for users without local compute resources. A 9B variant is currently in development and expected to be released soon.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.