Jul 2026· International Conference on Conversational User Interfaces· 0 citations· 55 references
Computer Science
TL;DR
While participants rated hedged and unhedged AI as equally trustworthy and likely to be correct, they were significantly less likely to follow hedged advice in a binary choice, and how linguistic markers can be used to calibrate user reliance to model certainty is discussed.
Abstract
As artificial intelligence (AI) systems are increasingly deployed for complex decision-making, calibrating user trust to prevent overreliance on overconfident AI remains a critical challenge. This paper investigates linguistic hedging – the use of tentative language to soften claims and indicate limited certainty – to communicate AI uncertainty, and its impact on user reliance. In a two-part investigation into decision-making in the financial domain, a formative study (N=36) explored strategies for eliciting hedged responses in accordance with model confidence. A confirmatory study (N=71) then measured actual behavioral reliance in a financial investment decision task, manipulating both AI confidence and the decision risk. Our findings reveal that while participants rated hedged and unhedged AI as equally trustworthy and likely to be correct, they were significantly less likely to follow hedged advice in a binary choice. We discuss how linguistic markers can be used to calibrate user reliance to model certainty, reducing overreliance while preserving trust in the system.
It is argued that calibration-aligned design (rather than trust maximization alone) should guide the development and assessment of high-stakes AI decision support, because reductions in reported trust do not consistently translate into commensurate changes in reliance behavior.
As AI systems increasingly support consequential human decisions, the confidence they display shapes whether users accept or override their recommendations. A common assumption is that well-calibrated model confidence will induce well-calibrated human reliance. This paper shows that the two are distinct. It introduces calibrated reliance as a team-level property of human–AI decision making and proposes Reliance Calibration Error (RCE) as a metric for quantifying the gap between displayed confidence and realized team accuracy. Using 35,670 human–AI interactions from the HAIID benchmark and complementary evidence from the GRACE benchmark, the analysis identifies a systematic calibration–reliance gap: even near-calibrated models can produce miscalibrated reliance once confidence is interpreted through human judgment. The evidence is consistent with self-confidence moderating advice uptake, but the observational design does not identify anchoring as the unique mechanism. Because this gap arises at the interface between model confidence and human action, the paper develops a human-aware confidence communication framework that remaps displayed confidence without changing the underlying predictor. On held-out HAIID data, plug-in observed-outcome RCE indicates that subgroup-aware remapping can substantially reduce display–outcome misalignment, while model-predicted simulations show that bounded/global policies improve team MSE more reliably than aggressive subgroup remapping. Diagnostics also show that unconstrained remapping creates substantial boundary mass and that model-predicted counterfactual RCE is sensitive to aggressive display shifts. Bounded variants preserve much of the simulated decision-quality gain while avoiding 0/1 displays. A small prospective pilot is reported as a feasibility check rather than confirmatory validation. These findings suggest that confidence interfaces should be evaluated for human–AI team behavior, with explicit support, transparency, and robustness checks, rather than for model-side statistical fidelity alone. Code and materials for reproducing the analyses and using the confidence-display policies are available at https://github.com/OliverDOU776/From-Calibrated-Confidence-to-Calibrated-Reliance.
Zijian Wang, K. Hu· ACM Transactions on Social C...· 0 citations
The findings support the AI aversion hypothesis and suggest that there needs to be more work to identify how best to explain AI-based calculators to foster trust and comfort more specifically.
Madhuri Ramasubramanian, Brian J. Zikmund-Fisher· Medical decision making· 0 citations
Analysis of consumers' trust in AI-generated recommendations under conditions of AI-assisted decision-making shows that emotional trust may be a mediator in the intention to delegate decision-making to AI agents and proposes strategies to build more transparent and trustworthy AI recommendation systems that can improve the user experience.
Yayi Liu· Frontiers in Humanities and...· 0 citations
This work develops a six step BBN framework and illustrates it to model customer intention to consult a doctor in an alternative healthcare system and reveals that while self efficacy appears to be a major factor, its actual causal impact is small.
Kumar Rahul, Shovan Chowdhury Indian Institute of Management Kozhikode, Kerala et al.· 0 citations
This study aims to address the growing concerns surrounding the use of artificial intelligence (AI) in property valuation, particularly issues of transparency, trust and accuracy. This study focuses on aligning AI models with expectations to foster responsible adoption in real estate decision-making.
A systematic literature review (SLR) of 44 peer-reviewed studies published between 2018 and 2025 was conducted. NVivo software was used for qualitative coding, and the political, economic, social, technological, legal and environmental framework guided the analysis of external factors influencing AI adoption. The study examined both technical model performance and stakeholder concerns.
Random forest and support vector machines were most frequently applied in structured valuation tasks, while artificial neural networks were reported in contexts involving non-linear modelling and complex data patterns. Despite demonstrated predictive capabilities, policymakers and professional stakeholders placed greater emphasis on transparency, explainability and legal accountability. Identified trust-related challenges included algorithmic bias, limited model interpretability, regulatory ambiguity and insufficient integration of contextual factors. These findings informed the development of a hybrid AI valuation framework that integrates technological performance with governance mechanisms, contextual calibration and professional judgement to strengthen property valuation quality.
This study is limited by the absence of primary data from stakeholder interviews. The findings are based solely on published literature, which may not fully capture real-time industry perspectives or emerging on-the-ground challenges.
The proposed framework offers a transparent, data-driven solution for valuers, investors and regulators, supporting better-informed decisions and encouraging ethical AI adoption in real estate.
This study synthesises technical and stakeholder dimensions of AI in property valuation using a structured qualitative approach via SLR. It proposes a novel hybrid framework that integrates stakeholder trust factors with model precision to enhance both reliability and acceptance of AI tools.
Wajhat Ali, D. Samarasinghe, Zhenan Feng et al.· Urbanization, Sustainability...· 0 citations