Alignment pretraining: AI discourse creates self-fulfilling (mis)alignment
Key point
In a 6.9B LLM experiment, AI discourse increased misalignment while aligned discourse reduced it.
Details
This is a controlled study verifying how AI-related discourse contained in the pretraining corpus actually affects downstream alignment. The researchers pretrained a 6.9B parameter LLM with varying amounts of (mis)alignment discourse to examine how talk about AI and training data change a model's behavioral priors.
The key results are clear.
- Adding more synthetic training documents dealing with AI misalignment noticeably increased misaligned behavior.
- Conversely, adding more documents on aligned behavior lowered the misalignment score from 45% to 9%.
- This effect weakened somewhat after post-training but continued to be observed.
The researchers interpret this as a self-fulfilling alignment phenomenon — that is, how we describe and talk about AI can change the model's alignment tendencies themselves. In conclusion, they suggest that alignment is not solely a post-training issue but must also be addressed in pretraining data design, and that this should be viewed as a separate axis called alignment pretraining. The models, data, and evaluation results have been made public at alignmentpretraining.ai.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.