Skip to content
Review Open access

Trust-aware evaluation frameworks for large language model reliability in enterprise AI platforms

Jul 2026 · International Journal of Science and Research Archive · Vol 20, pp. 889-898 · 0 citations

TL;DR

This article argues for an evidence-based approach that connects model behaviour, platform controls, user reliance, and auditable governance in enterprise reliability assessment in trust-aware LLM evaluation.

Abstract

Trust-aware evaluation is an emerging but rapidly consolidating research area for assessing the reliability of large language models in enterprise AI systems. Generative AI brings a new set of reliability challenges that extend beyond traditional accuracy metrics: An incorrect answer could affect key business decisions, customer interactions, knowledge management, security setups, or organizational accountability. This review summarizes peer-reviewed journal articles published between 2015 and 2025 related to key aspects of trust-aware LLM evaluation, including hallucination and factuality assessment, LLM evaluation methods, trustworthy AI governance, and human trust calibration. As indicated in the literature, current evaluation practice is still fragmented, both in terms of the various technical metrics and in terms of the documentation instruments, as well as on the interpretation of the data and the user-centred design of the trust cues, organizational governance. The most critical gaps are weak alignment between benchmark outcomes and enterprise risk, limited post-deployment monitoring, insufficient context-specific trust calibration, and limited validation of evaluation frameworks in operational platforms. The article argues for an evidence-based approach that connects model behaviour, platform controls, user reliance, and auditable governance in enterprise reliability assessment.

Read PDF

Similar papers

Review Open access Aug 2026

Requirements elicitation for public-private trust framework development

Public-private data sharing increasingly depends on frameworks that align participating organisations around shared definitions, responsibilities, rules, and agreements. Yet, organisations are struggling to design these trust frameworks for public-private data sharing. Based on 11 challenges derived from an LLM-assisted systematic literature review on Diffusion of Innovations Theory and Coordination Theory, we develop role-based requirements for a document-grounded conversational assistant intended to support the development and maintenance of public-private trust frameworks. We then prioritised and refined the challenges and requirements through an expert survey. The proposed assistant concept functions as a shared “knowledge navigator” over collaboration documents: it answers questions with traceable sources, surfaces knowledge gaps in the underlying corpus, guides onboarding and participation decisions, and supports agreement drafting and review by generating text while flagging inconsistent terminology as documents evolve. The outcome is threefold: a prioritised set of document-mediated challenges in public-private trust framework development; an expert-refined, role-based set of functional and non-functional requirements for a document-grounded conversational assistant; and a synthesis that places trust framework development along two dimensions, internal versus external and substance versus process.

Louise van der Peet, Nitesh Bharosa, M. Janssen · 0 citations
Open access Aug 2026

When Trust Sharpens Skepticism: Large Language Models in Indonesian Audit Practice

This study examines the effects of perceived transparency, explainability, and social influence on auditors' trust in Large Language Models, and the effect of that trust on their professional skepticism at Indonesian public accounting firms. The gap between global acceptance and trust levels toward artificial intelligence systems suggests that technology adoption is not always matched by adequate evaluation, a condition relevant to auditors, who must remain critical toward Large Language Models given their tendency to produce inaccurate answers. This study used a quantitative approach with Partial Least Squares Structural Equation Modeling, involving 102 auditors selected through purposive sampling. Results show that all three antecedent variables positively and significantly affect auditors' trust in Large Language Models, with perceived transparency contributing most, while trust in Large Language Models also positively and significantly affects professional skepticism, a direction opposite to the reliance pattern reported in prior audit automation literature. These findings suggest that trust in artificial intelligence based technology can form in a calibrated manner, coexisting with auditors' awareness of system limitations rather than diminishing their professional skepticism.

Kamal Amarullah, H. Ritchi, A. Mubarrok · 0 citations
Review Jul 2026

Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review

Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to understand. Developers struggle to assess the validity of LLM-generated reviews, making it difficult to gauge how much trust to place in them. The role of Explainable AI (XAI) in code review and its impact on trust remain underexplored. Objective: We study the influence of XAI on developer trust in AI-assisted code reviews. Method: We conducted a within-subjects user study with 34 participants, comparing three LLM-based code review systems with varying levels of XAI support: Condition A (detailed explanation and review feedback), Condition B (review feedback only), and Condition C (no explanations). Participants reviewed real-world code change requests alongside the AI-generated reviews. We measured trust perceptions, agreement with the AI recommendation, the reasoning given for each decision, and the time taken. Results: The level of explanation significantly influences both trust and agreement with AI recommendations, but in different ways. Full explanations (A) yield the highest perceived trust (M = 3.99/5) but not the highest agreement, whereas moderate explanations (B) achieve the highest agreement (89.22%). This could suggest that more explanation prompts developers to question AI recommendations more frequently. No explanations (C) results in the lowest trust and agreement. Explanation level did not significantly affect review time. The most commonly cited reasons for decisions were code readability and correctness. Conclusion: Incorporating XAI into code review significantly changes trust perceptions and agreement with AI recommendations. These results inform the design and evaluation of trustworthy AI-based code review systems, as well as studies on the human factors of AI-assisted software development.

Zhenhan Gao, Marvin Muñoz Barón, Umm E. Habiba et al. · 0 citations
Review Open access Aug 2026

Communicating Risk in the Age of Misinformation: Empirical Frameworks for Understanding Stakeholder Trust

ABSTRACT This systematic literature review examines how stakeholder trust is conceptualized and measured in risk and science communication amid increasingly complex information environments shaped by misinformation and disinformation. Following PRISMA procedures, which provide a framework for conducting a systematic review with transparency, replicability, and completeness, we searched major databases and screened studies using explicit inclusion criteria. For instance, trust had to be operationalized and measured in risk or scientifically complex contexts; credibility studies were retained when credibility was a trust‐related dimension related to stakeholder communication. Scientifically complex contexts are taken to refer to settings where information is inherently complex and nuanced, and understanding requires advanced knowledge. The final corpus (k = 69) includes quantitative, qualitative, and mixed‐methods designs, with a post‐2020 increase aligning with pandemic‐era scholarship. The findings identify three roles for trust. First, stakeholder trust as an outcome is strengthened more by transparency, empathy, and perceived value similarity than by expertise claims alone. Second, trust as a broad concept often embodies more than just detailed technical understanding; rather, trust predicts acceptance of technologies and policies and lower perceived risk. Third, as a mechanism, stakeholder trust mediates links between credibility, transparency, or value similarity and downstream outcomes such as compliance and cooperation. From organizations to interpersonal interactions, stakeholder trust formation reflects the characteristics of trustors, trustees, and context. We integrate these strands into a communication approach that emphasizes clarity about uncertainty, engagement, and visible accountability as preconditions for trust. Practically, the review underscores the design of transparent messaging, alignment with public values without oversimplification, and participatory approaches that treat stakeholders as partners. Conceptually, we separate judgments about whether information seems credible from trust in the people or institutions behind it. We argue that trust helps connect what people think and feel to what they are willing to do, especially when they must make decisions under uncertainty in today's complex information environment.

Matthew S. Weber, Chaeyeong Margo Lee, D. Kosson · 0 citations
Review Aug 2026

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic on January 27, 2025, when Nvidia lost USD589 billion in market value. Yet vendor benchmarks often depend on an honor system. Academic reassessments and independent leaderboards have found undisclosed changes to proprietary models, contaminated training data, and selective reporting. LLM-as-a-judge methods scale evaluation by reducing human review. Studies, however, suggest that judges may show identity-aware bias, scoring an answer according to its source model rather than its quality. This bias has not been fully measured or corrected across politically sensitive, reasoning-intensive, and preference-based tasks. We examine this problem using seven verifier models: GPT-OSS 120B, Llama 3.3 70B, GLM 5.1, Qwen3 32B, DeepSeek V4 Pro, Mistral Large3, and Sarvam M. They score anonymous and identity-disclosed responses from three primary models on 58 factual, reasoning, political, and preference-based questions. Identity disclosure slightly raises scores for factual questions, moderately affects stress-reasoning tasks, and causes large changes for geopolitically sensitive topics. Notable results include GLM5.1 (+7.00 points, p = 0.0249) and Llama 3.3 70B (+1.56 points, p = 0.00). We also introduce a blockchain-based commit-reveal protocol using Autonomous Economic Agents on an Ethereum-compatible ledger. In Phase 1, each judge records a one-way hash of its score and a secret salt before candidate identities are revealed. In Phase 2, the identity and raw score are disclosed and verified on-chain. This creates a tamper-evident audit trail that separates blind evaluation from post-hoc claims and reduces the verification burden on independent researchers and leaderboard operators.

Sahil Pardasani, Madhusudan Singh · 0 citations
Review Aug 2026

DCI: Dependency Confidence Index for Assessing Open-Source Dependency Trustworthiness

Selecting trustworthy open source software dependencies remains a major challenge in software supply chain security. We present the Dependency Confidence Index (DCI), a composite formative index that combines nine empirically weighted trust factors into a single normalized composite score for dependency selection. DCI's trust factors combine insights from a systematic literature review and an exploratory Analytic Hierarchy Process (AHP) survey of ten software developers, highlighting security, source code quality, and project health as the most influential dimensions. Following Goal-Question-Metric methodology, we implemented 12 automated measurements using SonarQube, GitHub APIs, and OpenSSF Scorecard data, deployed in a containerized evaluation platform. We conducted a pilot evaluation of the normalized DCI on 92 popular PyPI packages, observing moderate agreement with OpenSSF Scorecard scores and perfect test--retest reliability. Analysis reveals process-based factors (dependency management, CI) dominate scores on high-quality packages, while security metrics saturate---suggesting DCI's complementary role to existing tools. Our publicly available implementation provides a foundation for open source software trustworthiness research and practical dependency auditing.

C. Albrecht, Stefan Reitmann · 0 citations