Limits of LLMs in Simulating Human Preferences Confirmed
Key point
LLMs predicted actual human choices with only a 53% probability, revealing their limitations as a simulation tool.
Details
Recently, there has been a trend of companies trying to use LLM-based 'Synthetic Users' instead of real users to cut costs, but a new study presents skeptical data on this.
When researchers tested LLMs using 28 real study cases and 78 choice tasks, the probability that LLMs matched the majority opinion of humans was only 53%. This is no better than a coin flip for two-option choice tasks.
The key findings of the study are as follows:
- Limits of Personas and CoT: Assigning detailed personas or applying Chain-of-Thought (CoT) reasoning produced almost no improvement in performance.
- Homogenization of Justifications: Rather than reflecting the diverse experiences of real humans, the model's reasoning process instead homogenized the outputs, resulting in lower semantic similarity to the actual grounds for human judgments.
- Nature of Training Data: This is analyzed to be because LLMs are not designed to predict human preferences, but rather trained to reproduce outputs that humans would be likely to prefer.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.