Skip to content
#generative ai Preprint

AffAdapt: AFFect-driven ADAPTive AI Personas for Seamless Conversations

Aug 2026 · 0 citations · 12 references
Computer Science

TL;DR

AffAdapt is presented, a seamless interaction design framework for AI-personas, which coordinates streaming speech recognition, proactive turn-management, persona-grounded response generation, a persistent emotional state, and synchronized embodied output into a single interaction loop.

Abstract

AI-generated personas are being increasingly used for support, training and simulations. While generative AI models possess abilities to generate affect-aware responses, their embodiment into visual personas is an active area of investigation. Naturalistic exchanges require understanding of the conversational partners'turn completions, whether the agent should respond or keep listening and rely on non-verbal cues aligned with one's emotional states. Seamless human-AI conversation in a multimodal setting requires all modalities being generated to act in coordination. We present AffAdapt, a seamless interaction design framework for AI-personas, which coordinates streaming speech recognition, proactive turn-management, persona-grounded response generation, a persistent emotional state, and synchronized embodied output into a single interaction loop. We demonstrate the architecture in the context of practicing sensitive, high-stakes conversations, and report an initial case study showing fluid turn management and adaptive, persona-consistent behavior, alongside open challenges in interruption handling, open-ended dialogue, and multimodal affective alignment. AffAdapt's interaction loop is a generalizable pattern for coordinating timing, identity, and affect in real-time AI personas - applicable to training, coaching, education, and simulation contexts wherever believable, responsive interaction matters.

View source

Similar papers

Book Open access Jul 2026

Emerging Risks of Conversational AI

Conversational User Interfaces (CUIs) are rapidly advancing, moving from single-task assistants to powerful and engaging artificial agents. As CUIs become embedded in daily life, the implications for users interacting with such systems demand closer scrutiny. While some issues are immediately identifiable (e.g. privacy risks, misinformation, AI hallucinations), others may only become visible after longer periods of use, including over-reliance, erosion of human agency, or the normalization of biased or exclusionary language. This workshop aims to gather a multidisciplinary community to reflect on pressing challenges and uncertainties and explore risk analysis and mitigation strategies in CUI design and deployment. Through presentations, discussions, and collaborative design activities, participants will examine immediate and longitudinal risks and share empirical insights. The activities will support developing interdisciplinary frameworks to understand and handle unintended consequences of CUIs. The workshop seeks to build an international network of scholars and practitioners to promote responsible and human-centred conversational AI.

Manveer Kalirai, C. Wei, Thomas Essmeyer et al. · 0 citations
Book Open access Jul 2026

ADAPTIC: Adapting Dialog and Pragmatic Traits in Context

The rise of LLMs has enabled CUIs to increasingly mimic human social and conversational cues, e.g., tone of voice and emotional expressions. However, this mimicry usually lacks strategic communicative intent, placing the cognitive burden of mutual understanding on the user. At the same time, CUIs based on general-purpose, task-agnostic LLMs are being deployed across varied domains with distinct, context-specific conversational needs, including, e.g., healthcare, education, and journalism. Therefore, there is a growing need to transition from arbitrary, domain-agnostic generation of pragmatic cues to strategic adaptation of both visual and linguistic interface features. This workshop proposes a paradigm shift toward designing context-specific CUIs that actively support communicative success through pragmatic cues. Bringing together perspectives from HCI and social sciences, we will explore how users appropriate conversational AI across domains. Through cross-disciplinary dialogue, the workshop aims to establish a shared vocabulary, identify domain-specific challenges, and lay the groundwork for future collaboration.

Laura Spillner, Johanna Rockstroh, Paul Goerke et al. · 0 citations
Preprint Jul 2026

PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction

Equipping humanoid robots with coherent and adaptable personas is crucial for fostering natural, engaging, and trustworthy human-robot interaction (HRI). However, existing approaches often rely on static, hard-coded identities that lack the flexibility to adapt to individual user contexts. In this paper, we present PACE (Persona Adaptation through Conversational Elicitation), a novel framework for the interactive generation and deployment of structured personas on the Ameca humanoid robot. Our system introduces an Interactive Persona Elicitation Pipeline, enabling the robot to dynamically synthesize a tailored, psychologically grounded identity through user Q&A. This elicitation process feeds into a persona prompt compilation phase, generating a structured persona prompt built upon multi-perspective dimensions. We detail the Embodied System Integration required to translate this structured specification into expressive, multimodal humanoid behaviors. Through a comprehensive empirical HRI evaluation, we assess the impact of dynamically generated personas on user trust, perceived anthropomorphism, persona consistency, personal relevance, and interaction quality compared to a generic baseline. These contributions establish a scalable pathway for deploying personalized, interactive, and reliable identities in embodied humanoid assistants. Video demo is available at: https://lipzh5.github.io/PACE/

Peizhen Li, Longbing Cao, Megani Rajendran et al. · 0 citations
Preprint Aug 2026

Aura: Dynamic Intra-Turn Emotion-Aware Adaptation of Large Language Model Responses

Effective human-AI interaction requires systems that dynamically adapt to a user's behavior and evolving understanding. When users interact with Large Language Models (LLMs), these models typically respond to prompts without sensing the user's immediate reactions. This lack of communicative synchrony can lead to information overload or leave confusion unresolved in real time. In this paper, we introduce Aura, a framework that enables LLM systems to dynamically modulate output based on a user's evolving emotions. Aura's Perception Module continuously estimates the user's emotional state from facial expressions. Our Policy Module then selects interventions through a probabilistic belief model. Finally, Aura's Generation Module uses parameter-efficient Low-Rank Adaptation (LoRA) adapters to produce contextually tailored responses mid-turn during response generation. We evaluated Aura in a within-subjects user study (N=20) on information-seeking tasks, where it achieved statistically significantly higher normalized perceived learning gains than a Llama-3 baseline and reduced interaction time by 21% relative to existing LLM baselines (GPT-4o, Llama-3). Our results indicate that real-time, context-sensitive interventions can improve learning efficiency and user satisfaction without observable degradation in factual accuracy. Aura thus supports the potential for more responsive and effective human-AI interaction.

Rachel Schuchert, Christian Holz · 0 citations
Review Jul 2026

HARP: The Human--AI Research Platform

Large language models (LLMs) have shifted human--computer interaction from `traditional''interface journeys toward more conversational exchanges. Researchers studying HCI and UI use moderated usability sessions, interviews, surveys, transcript analysis, and static prototypes. However, static prototypes provide limited opportunities to study interaction with live AI systems or systematically control how an LLM behaves across participants and scenarios. Conversation transcripts reveal little about how users formulate, revise, and hesitate over prompts before submission. We designed the Human--AI Research Platform (HARP) for researchers, designers, and anyone who has ever wondered, `What if AI did this?'HARP places participants in controlled mock scenarios with live, configurable AI agents. Researchers can control agent prompts, model parameters, response characteristics, and experimental conditions; trigger surveys at predefined moments; and record prompt composition time, response latency, deletions, and keystroke pauses. Planned capabilities include voice, facial expression, gesture, and, where legally and ethically appropriate, emotion analysis. We illustrate HARP through a study examining how technical specificity and response length affect retention of LLM output. By pairing controllable live agents with behavioral and self-report measures, HARP enables systematic testing of how AI design choices affect users.

Zeshu Zhu, Natalie Friedman, Kevin Weatherwax et al. · 0 citations
Conference Jul 2026

A Multimodal Emotional Interaction Framework Driven by Large Language Models

Natural and empathetic human-robot interaction is essential for social robots and other AI applications, while emotional feedback is less discussed. Thus, this paper proposes a multimodal interaction system driven by large language model (LLM). The system constructs a unified emotional state vector by integrating visual (facial expressions) and auditory (speech emotion) cues. It employs DeepSeek LLM for context-aware chain-of-thought reasoning to generate contextually appropriate verbal responses, facial expressions, and head movement commands. To achieve optimized latency and fluid embodied interaction, the system adopts a layered architecture: the upper-level LLM handles semantic understanding and behavior planning, outputting structured JSON commands; the lower-level controller generates smooth motion trajectories and manages multimodal interaction flows using a finite state machine (FSM). Experimental results demonstrate the system’s ability to effectively resolve emotional ambiguities, track emotional evolution during continuous dialogue, and achieve an optimized end-to-end response latency. This validates its feasibility and engineering value in practical human-robot interaction scenarios.

Yuxuan Chen, Chen-Yi Qiu, Ning-Chuan Wang et al. · 0 citations

Related blog posts