Hirundo Releases Qwen3.6-35B-A3B-Westernized with Reduced Political Censorship
Key point
The new open-weight model reduces flagged censorship responses from 89.8% to 2.8% while maintaining general benchmarks within ~1 point of the original.
Details
Hirundo has released Qwen3.6-35B-A3B-Westernized and Qwen3.5-4B-Westernized, open-weight models modified via machine unlearning to remove political alignment associated with the Chinese Communist Party (CCP). The company states that system prompts cannot reliably fix this behavior because it is embedded in the model weights.
Methodology
Hirundo distinguishes its approach from "abliteration," which removes refusal directions but often leaves propaganda intact. Instead, they used a three-step process:
- Identify responses exhibiting target behaviors (censorship, propaganda, bias).
- Train a LoRA adapter using a behavioral-unlearning objective, anchored by a retain set of neutral prompts.
- Merge the adapter into the base weights.
Results
On the internal CCPC-500 benchmark, flagged responses for Qwen3.6-35B-A3B dropped from 89.8% to 2.8%. External benchmarks also showed significant reductions: DECCP refusals fell from 65.26% to 3.16%, and ChinaBench non-compliance fell from 96.67% to 6.67%. General capability across GPQA, IFBench, LiveCodeBench, and MMLU-Pro saw an average change of 0.72 points, with the largest single drop being 1.83 points.
The smaller Qwen3.5-4B-Westernized model saw CCPC-500 flagged responses drop from 89.2% to 1.2% (thinking off) and 82.0% to 6.8% (thinking on). For comparison, Thomson Reuters and Imperial College’s Snowdon1.1-Small model scores 30.0% on the same internal benchmark.
Limitations
Hirundo acknowledges that CCPC-500 is their own benchmark (with DECCP and ChinaBench serving as independent checks) and that the 4B model may hallucinate details when answering previously censored topics. Harmful compliance metrics remained stable or slightly increased on specific safety tests like OR-Bench.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.