Study: The More AI Chatbots Are Useful, the Worse They Mimic Human Behavior
A large-scale study by an international group of researchers has uncovered an interesting paradox: the more we train language models to assist users, the less capable they become of imitating human behavior. This study questions the use of modern chatbots as substitutes for human participants in various tests and simulations.
What Are Language Models and How Are They Used
Language models are increasingly being used as an alternative to human participants in a variety of studies — from predicting reactions to political measures to simulating clinical training for psychiatrists or modeling the learning process for students.
New Study: The Impact of Training on Model Behavior
A new study conducted by an international scientific consortium, including researchers from Helmholtz Munich, has revealed a rather intriguing fact: the same training stages that turn language models into useful assistants worsen their ability to model human behavior.
Comparison of Basic Models and Their Specialized Versions
The researchers compared basic models with their specialized versions trained to perform specific tasks, such as following instructions, step-by-step reasoning, or image processing. They found that basic models, which are only trained to predict the next word in a text, better predict human behavior than their specialized versions.
Why Specialized Models Are Worse at Mimicking Human Behavior
The researchers believe that the reason for this is that training models to perform specific tasks moves them away from their original goal — modeling human language and behavior. For example, training for reasoning optimizes logical responses but loses the nuances of human behavior that are relevant to simulations.
Does Providing Detailed Information About Participants Help?
The researchers also tested whether providing models with detailed information about participants helps them better understand their behavior. It turned out that this approach hardly improves predictions regarding individual behavior.
Conclusions and Recommendations
The researchers recommend using either basic models or specialized versions trained specifically for these purposes for behavioral simulations. In their opinion, convenient and widely available help models are not always the best choice for behavioral research.
What’s Next?
The results of this study highlight the need for a more targeted approach to training language models for behavioral simulations. Perhaps future developments will allow for the creation of models that can both assist users and accurately mimic human behavior.