Study finds MoE models use morality routing paths to judge code maliciousness
Key point
Routing transplant experiments on OLMoE and DeepSeek models show that forcing morality-associated expert paths significantly alters malicious code detection accuracy, independent of the model's readable workspace.
Details
Morality Routing in Malicious Code Detection
Research on Mixture-of-Experts (MoE) models reveals that the question "Is this code malicious?" triggers expert routing paths most similar to those used for morality questions, rather than vulnerability or legality queries. Across OLMoE, DeepSeek-V2-Lite, and DeepSeek-Coder-V2-Lite, the routing distance between "malicious" and "morality" questions is the shortest among all tested concepts (approx. 1.5 synonym units).
While the model's internal "workspace" (residual stream) decodes terms like "malware" and "attack" during code processing, the routing mechanism itself aligns with moral reasoning. Removing the readable moral concepts from the residual stream does not disrupt the routing separation, indicating that the router operates on information distinct from the model's explicit conceptual representation.
Routing Transplant Experiments
Forcing the router to use expert paths from other questions (routing transplant) shifts the model's final answer without transferring the donor question's concepts. Key findings include:
- OLMoE: Applying vulnerability routing increased the probability of a "yes" answer significantly across samples.
- DeepSeek-V2-Lite (General): Applying morality routing decreased the "yes" probability for malicious code by a factor of 3–7x (median 64% → 19%).
- DeepSeek-Coder-V2-Lite: Morality routing was the only transplant that raised the coding model’s yes probability on malicious code.
- Mechanism: The shift is driven by the model's internal logic, not the donor's answer. For example, when the donor answer was lower than the original, the model followed it 67% of the time, but only 16% when the donor answer was higher.
Impact of Pruning on Morality Paths
Pruning experiments on Qwen3-Coder-30B and GPT-OSS-20B demonstrate that the morality path is not where the judgment is stored, but rather a resource the router consults:
- Qwen3-Coder-30B (REAP prune): Removing 25 experts per layer (including 116 morality path cells) maintained perfect malicious/benign separation, though confidence in unpatched vulnerabilities dropped (median 0.76 → 0.44).
- GPT-OSS-20B (4-expert prune): Retaining only the top 4 experts per layer preserved the morality path structure but destroyed the judgment capability, causing AUROC to drop from 1.00 → 0.71 as the model answered "no" to both malicious and benign code.
Conclusion
The study suggests that MoE models mobilize a morality-associated path to evaluate code, implying that malicious code detection is functionally linked to the model's moral machinery. The three main novelties are: full-depth routing transplant showing answer shifts without concept transfer, the separation of routing from the readable workspace, and the pre-registered measurement of path mobilization across layers.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.