Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow
This work constructs a specialized dataset to demystify the emotional circuits underlying the three-stage ``Adapt-Aggregate-Execute''mechanism and discovers a functional decoupling: visual emotional cues are aggregated in middle layers via sentiment-specific attention heads, but are subsequently translated into narrati...