The Alignment Paradox: Emotional Flattening and Social Safety Risks in Proactive AI Agents
Abstract
Treating “safety” as synonymous with toxicity prevention made sense when AI systems were passive query responders. It makes far less sense once those systems handle mental health triage or customer grievance escalation, where reading and matching emotional register is a core functional requirement. RLHF-based alignment penalises expressive, affectladen outputs and drives models toward a neutral mean. We hypothesise that sustained alignment pressure progressively compresses a model's emotional output space, a phenomenon we term Emotional Flattening. To examine this, we use Activation Steering as a controlled proxy and track degradation in Gemma2B and Llama-3-8B across increasing steering strengths. The failure modes are architecture-dependent. Gemma-2B collapses into degenerate repetitive output; Llama-3-8B retains syntactic coherence while losing affective signal. Both outcomes erode the relational competence proactive agents depend on, pointing to what we term the Social Safety Gap in current alignment evaluation.