AI Briefing
KO

Show HN: We confirmed that distilling DeepSeek into GPT-OSS does not transfer censorship - try it yourself

·2026.07.31 03:13

Key point

A model trained on data from a Chinese model acquires the intelligence but does not inherit its political censorship.

Details

In a distillation experiment using outputs from a Chinese frontier model like DeepSeek to train a US model (GPT-OSS-120B), the model's financial reasoning ability improved significantly, but the political censorship tendencies specific to China did not transfer.

The key experimental results are as follows:

  • No censorship transfer: DeepSeek V4 Flash showed high censorship responses on sensitive topics related to China, but the GPT-OSS model trained on it answered those topics without censorship.
  • Performance improvement: On financial reasoning tasks, the distilled model effectively acquired the teacher model's intelligence.
  • Self-distillation effect: It was confirmed that similar performance gains could be achieved through self-distillation alone, without using a higher-tier Chinese model.

The research team has released the LineageEval evaluation tool used in this experiment, 304 prompt pairs, evaluation code, and model weights, all publicly available.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.