Qwen Comparative Analysis
Key point
A benchmark, safety, and tensor analysis comparing Qwen 3/3.5 with HauhauCS, Heretic, and Huihui was released.
Details
Three abliteration techniques (Heretic, HauhauCS Aggressive, Huihui) were applied to 5 Qwen models and compared.
The targets are Qwen3.5-2B/4B/9B/27B and Qwen3-4B-Instruct-2507; the Qwen3.5 series uses a hybrid Mamba2+Transformer architecture while Qwen3-4B uses a pure Transformer architecture, so differences in abliteration impact are examined as well.
Evaluation setup:
- Capability:
lm-evaluation-harness+vLLM, 8 tasks, bfloat16 - Safety: HarmBench 400 textual behaviours,
max_tokens=2048 - Additional analysis: weight analysis, KL divergence
The Qwen models were aligned as lossless safetensor files reverse-converted from BF16/FP16 GGUF for comparison, and the full benchmarks along with reproduction materials were released in a HuggingFace collection.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.