Effects of Shot Size on Social Impressions in Humans and Multimodal AI
The same person can be judged differently depending on how tightly they are framed. As image- and video-based evaluations become increasingly common in settings such as hiring, screening, and interviews, this raises a practical question: does shot size change social-impression ratings in human observers and a multimodal large language model (GPT-4o)? From one master photograph of each of eight adults with neutral expressions, we created medium-shot (MS), close-up (CU), and extreme-close-up (ECU) versions while holding expression, pose, lighting, perspective, and output size constant. Sixty participants rated all eight identities in a balanced design, and a fixed GPT-4o configuration evaluated each of the 24 images in ten stateless repetitions. Both evaluators rated the same five outcomes: Trust, Competence, Likeability, Discomfort, and Approachability. In human crossed linear mixed models, tighter framing increased Discomfort and decreased the other four outcomes; all five MS–ECU contrasts remained significant after Holm correction. Discomfort showed the largest human MS–ECU change (b = +0.600), whereas Competence showed the smallest (b = −0.221) and decreased in 5 of 8 identities. GPT-4o showed the same overall direction of change across all five outcomes, and all five identity-level MS–ECU sign-flip tests remained significant after Holm correction. The predicted larger CU–ECU change was not supported in humans, and no GPT-4o outcome showed a significant transition difference. For Discomfort, the larger observed change occurred from MS to CU in humans but from CU to ECU in GPT-4o. Across both evaluators, tighter framing produced less favorable social impressions and greater Discomfort for the same neutral identities.