This work investigates whether personality-aware fine-tuning can reduce the consistency and controllability of personality-conditioned dialogue generation compared with instruction prompting alone, and indicates that fine-tuned models are not better at role-playing different personalities than their respective baseline models.
Abstract
LLMs are increasingly used in social simulations for socially interactive agents and robots, offering more flexibility than rule-based systems. However, even though they mimic human behaviour very well, there is a persistent alienness to them. This work investigates whether personality-aware fine-tuning can reduce this gap by improving the consistency and controllability of personality-conditioned dialogue generation compared with instruction prompting alone. We fine-tune two small open-weight LLMs, Qwen2.5-7B-Instruct and Ministral-8B-Instruct, using a corpus that combines personality-labelled social media posts and dialogues to create a personality-based dialogue engine for social simulation. The resulting models are evaluated across multiple social interaction scenarios using three independent LLM judges, which assess personality fidelity and provide evidence-based behavioral interpretations. We additionally quantify inter-rater agreement and lexical characteristics of the generated dialogue. Results indicate that fine-tuned models are not better at role-playing different personalities than their respective baseline models. However, low inter-rater agreement limits the confidence with which these results can be interpreted. Concerning the quality of generated texts, fine-tuned models are mostly comparable to the baselines, with fine-tuning improving the linguistic diversity of the Qwen models. While the results appear generally usable and the baseline models offer the best overall performance, future studies should place greater emphasis on the quality and domain alignment of training data for accurate personality role-playing.
LLM-based conversational agents generate fluent responses but remain limited in adapting their supportive style to individual personality and emotional needs. We present a Detect–Regulate–Evaluate (D–R–E) architecture that performs turn-by-turn Big Five detection and applies Zurich Model-inspired behavioural regulation...
Duojie Jiahua, Samuel Devdas, Mirjam Stieger et al.· Electronics· 0 citations
Large language models are increasingly used to simulate human participants in social and behavioral studies, yet static persona prompting typically maps a participant profile and an experimental scenario directly to a response, entangling stable dispositions with situation-specific interpretations. To address this limi...
Zhi-Bo Deng, Dong-Yuan Li, Shu-Wen Ge et al.· 0 citations
Personality plays a central role in human-robot interaction, shaping how people engage with social robots. Yet most existing systems treat personality as fixed or reduce it to binary categories, limiting adaptation across diverse users. Moreover, adaptation is often implemented at the level of isolated components, rath...
Antonio Andriella, Giuseppina Russo, Silvia Rossi· IEEE Robotics and Automation...· 0 citations
Reinforcement fine-tuning (RFT) is increasingly used in applications where large language models (LLMs) interact with humans and other agents. Here we use social deduction games to study how RFT changes LLMs'social behaviour. We let fine-tuned and base LLM agents play hidden-role games that require hidden-state inferen...
Ling-Zhe Zhang, Yun-Peng Zhai, Tong Jia et al.· 0 citations
This work proposes A-B-D to infer traits bottom-up from behavioral data (B-data), namely how agents act on their environment and communicate with users, as recorded in existing trajectories, and offers a new lens for understanding AI personality.
Hao-Kai Zhao, Jie Gao, Yunze Xiao et al.· 0 citations