AI Briefing
KO

UkisAI Releases Swift Model Boosting Qwen 3.8 27B Inference Speed by 1.95x

·2026.09.15 00:57

Key point

UkisAI released the Swift model, which reduces unnecessary inference tokens in Qwen 3.8 27B by 58% and improves speed by 1.95x.

Details

UkisAI released the Swift-Qwen3.8-27B model, which maximizes inference efficiency based on Qwen 3.8 27B. By identifying and optimizing overthinking patterns, this model reduces thinking tokens by 58% and improves speed by 1.95x. Accuracy loss remains under 1% on most benchmarks.

Optimization Method and Results

Unlike existing methods that forcibly shorten inference length, this approach applies penalties only to 'anxiety-like' inference loops discovered in PTQ and BF16 models. By generating diverse domain traces in an 8xH100 environment to identify common tokens, accuracy was restored through LoRa SFT and On-Policy Distillation.

Key benchmark comparison results:

  • GPQA-Diamond: Accuracy 88.4% → 88.3% (tokens reduced by 58%)
  • LiveCodeBench v6: Accuracy 76.8% → 81.6% (tokens reduced by 46%; note that this figure is due to baseline truncation and does not represent pure performance improvement)
  • Terminal-Bench 2.1: Accuracy 66.7% → 65.8% (tokens reduced by 39%)
  • MMLU-Pro: Accuracy 85.5% → 85.0% (tokens reduced by 28%)

Deployment and Limitations

The model is released on Hugging Face, offering various quantized versions such as GGUF and NVFP4, along with an OpenAI-compatible API. However, on the AIME 2026 benchmark, a 4.6% accuracy loss occurred due to a specific token penalty bug, which is scheduled to be fixed in future updates. This method is a complement, not a replacement, for reasoning effort settings, aiming to remove only unnecessary thinking parts while maintaining xhigh-level accuracy.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.