Skip to content
Preprint

Psychological Competence as a Missing Dimension in AI Evaluation

Jul 2026 · 0 citations · 43 references
Computer Science

TL;DR

It is argued that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.

Abstract

Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The relevant unit of evaluation is therefore not only the model, but the human-AI interaction. This paper introduces psychological competence as a missing dimension in AI evaluation. We define psychological competence as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways that are appropriate to the user, context, and purpose of the interaction. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluation approaches capture parts of this problem but rarely assess these psychological effects directly. Drawing on behavioral science and human-AI interaction research, we outline a conceptual framework for psychological competence and its core domains. Rather than proposing a specific benchmark, we define the construct, clarify its boundaries, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. We argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.

View source

Similar papers

Review Open access Jul 2026

Evaluating AI “Understanding” with Cognitive-Psychological Criteria: Evidence, Gaps, and Structural Limits

This article clarifies the concept definitions and evaluation criteria of understanding in cognitive psychology by combining classic theories and experimental evidence, and uses these criteria as the analytical framework for the performance of “similar understanding” in contemporary artificial intelligence systems.

Ruo Qin · 0 citations
Review Open access Aug 2026

Artificial Intelligence in eWOM: Roles, Mechanisms, and an Integrative Framework

An AI-enabled lifecycle of Creation, Transformation, Transmission, Evaluation, Evaluation, and Governance is proposed and an AI-eWOM fit perspective is developed and a TCCM-organized research agenda identifies priorities for future research.

A. Joyal · 0 citations
Review Open access Aug 2026

Cognitive Readiness for Human-AI Collaboration.

ObjectiveThis narrative review examines the cognitive, metacognitive, and team competency requirements that may contribute to productive and reliable collaboration between human and AI to address two questions: What capabilities make AI a competent collaborator? What makes humans ready for AI collaboration?BackgroundAs AI systems are increasingly integrated into workplaces and framed as teammates rather than tools, humans face challenges that include maintaining situation awareness, calibrating trust, and working with systems that may surpass them cognitively. We analyzed Human-Agent Teaming (HAT) readiness around two complementary levels: operational team competencies (communication, coordination, and adaptability) and regulatory capacities (trust calibration and metacognitive awareness).MethodWe conducted a structured narrative review of literature from 2010 through January 2026, searching Google Scholar, Scopus, PsycINFO, IEEE Xplore, ACM Digital Library, and Semantic Scholar, complemented by forward citation tracking. After screening 572 records, 192 articles were included for synthesis.ResultsCommunication inflexibility, limited shared understanding, and trust miscalibration emerge as recurring barriers to HAT, while regulatory capacities (trust calibration and metacognitive awareness) represent particularly critical dimensions of HAT readiness that remain to be fully operationalized.ConclusionHAT requires mutual readiness, with both humans and AI developing metacognitive and adaptive capabilities. Despite methodological heterogeneity limiting clear conclusions, cross-training and co-learning methods offer a promising avenue for building shared understanding and calibrated collaboration.ApplicationThis review provides practical principles for designing AI systems that support calibrated collaboration and for preparing humans to work adaptively with AI, thereby enhancing team effectiveness, reliability, and resilience in collaborative work environments.

Sébastien Tremblay, Delphine De Hemptinne, Gabrielle Teyssier-Roberge et al. · 0 citations
Review Jul 2026

Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

It is argued that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications, and that addressing it is important for creating AI systems on which users are both willing and justified to rely.

Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins et al. · 0 citations
Book Open access Jul 2026

Intelligence Is Not All You Need: A Case for Artificial Wisdom in Conversational Agents

It is argued that intelligence alone cannot determine appropriate goals, guide action under uncertainty, or ensure beneficial human outcomes, and proposed artificial wisdom as a needed corrective for conversational systems.

Matthias Kraus · 0 citations
Preprint Jul 2026

Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles

Interaction Readiness is introduced as a framework for specifying and evaluating that missing layer of performance in role-bearing AI agents, and its findings are translated into a specification template and audit procedures that product and engineering teams can apply before and after deployment.

Sudhir Alladi Venkatesh · 0 citations