Multidimensional Analysis of Human-like Behavior in LLMs: Models, User Factors, and System Prompts
Key point
Analysis of 21,000 conversations across 4 LLM models verified the appropriateness and controllability of human-like behaviors.
Details
Large language models (LLMs) exhibit various human-like behaviors, such as expressing thoughts and emotions or forming relationships with users. However, researchers lacked systematic methodologies and empirical insights to determine when and in what forms these behaviors appear.
The research team analyzed 21,000 multi-turn conversations from 4 models: gpt-4o, gpt-4.1-mini, claude-sonnet-4.6, and gemini-2.5-flash. Using LLM-as-a-judge and human evaluation, they multidimensionally assessed the frequency, potential impact, and controllability of these behaviors.
The analysis revealed that while human-like behaviors are universal, they vary depending on model and user factors (conversation goals, user profiles). Human evaluators judged self-referential and relationship-building behaviors as less appropriate in LLMs than in humans, whereas boundary-maintaining behaviors were rated as more appropriate in LLMs.
The study also confirmed that these behaviors can be controlled via system prompts. It emphasized the need for careful evaluation to avoid unintended side effects and provided recommendations for responsible LLM design and evaluation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.