Skip to content

Category

large language models

451 papers

#large language models Open access Sep 2026

A Transformation-Based AI Framework for Equitable RTI Tier Classification in Qatar

Consistent and equitable classification of students within Response to Intervention (RTI) frameworks remains a significant challenge in Qatar’s educational system, where tier-assignment decisions rely predominantly on qualitative, multidisciplinary evaluation reports rather than standardized quantitative measures. This research proposes a transformation-based framework for RTI tier classification using Qatari student evaluation reports. The framework converts qualitative clinical descriptors into structured intermediate representations through descriptor extraction, ordinal severity mapping, score translation, and composite aggregation. Seven classification approaches were systematically evaluated within this framework: direct zero-shot and few-shot large language model (LLM) classification; hierarchical prompting; rule-based transformation; LLM-assisted transformation; and a hybrid transformation-based approach. All experiments utilized OCR-extracted Arabic evaluation report text and were assessed using accuracy, balanced accuracy, macro F1-score, weighted F1-score, and class-wise F1-scores. The hybrid transformation-based approach demonstrated the strongest overall performance across evaluation metrics and was the only approach to maintain meaningful classification performance across all three RTI tiers, including the underrepresented and most challenging Tier 1 category. Direct and hierarchical prompting approaches produced lower and less consistent classification performance across the evaluated RTI tiers. These findings indicate that the introduction of structured intermediate transformation stages substantially enhances the consistency, interpretability, and equity of RTI tier classification from qualitative evaluation reports, providing a principled mechanism for standardizing classification decisions across Qatar’s schools and evaluation teams.

Ali M. Alodat, Shadi Banitaan, Ahmad Aljaafreh et al. · 0 citations
#large language models Open access Sep 2026

Integrating LLM in Business Process Management: A Conceptual Framework for Augmenting the Process Lifecycle

Large language models (LLMs) are increasingly relevant to Business Process Management (BPM), particularly when process knowledge is dispersed across documents, conversations, and other unstructured sources. Their probabilistic outputs, however, raise questions about validation, traceability, and accountability. This paper develops a lifecycle-based conceptual framework for allocating and governing LLM use across the six stages of the BPM lifecycle. The framework separates generative interpretation from formal, empirical, and expert validation. It comprises five interdependent layers and six operational principles, implemented through a stage-risk-validation matrix, a principle-intensity map, and four evaluation dimensions. Governance requirements increase as outputs approach live execution or decisions that are difficult to reverse, with controls aligned with the NIST AI Risk Management Framework, the EU AI Act, and the GDPR. A customer complaint-handling scenario demonstrates how the framework can be applied. An illustrative stress test using the BPI Challenge 2017 event log and ten independent LLM generations instantiates the validation layer under information-asymmetric conditions. Although all generated models were structurally valid, the event log revealed incomplete activity coverage and control-flow mismatch. This illustrates the value of an external referent but does not establish comparative performance or a general difference in error detectability between LLM-generated and process-mining artefacts. The framework therefore positions LLMs as tools for turning unstructured information into preliminary process knowledge, while established BPM methods and human expertise remain responsible for validating consequential outputs.

Florin Dumitriu, Valerică Greavu-Şerban, Sabina-Cristiana Necula et al. · 0 citations
#large language models Open access Sep 2026

Effects of Shot Size on Social Impressions in Humans and Multimodal AI

The same person can be judged differently depending on how tightly they are framed. As image- and video-based evaluations become increasingly common in settings such as hiring, screening, and interviews, this raises a practical question: does shot size change social-impression ratings in human observers and a multimodal large language model (GPT-4o)? From one master photograph of each of eight adults with neutral expressions, we created medium-shot (MS), close-up (CU), and extreme-close-up (ECU) versions while holding expression, pose, lighting, perspective, and output size constant. Sixty participants rated all eight identities in a balanced design, and a fixed GPT-4o configuration evaluated each of the 24 images in ten stateless repetitions. Both evaluators rated the same five outcomes: Trust, Competence, Likeability, Discomfort, and Approachability. In human crossed linear mixed models, tighter framing increased Discomfort and decreased the other four outcomes; all five MS–ECU contrasts remained significant after Holm correction. Discomfort showed the largest human MS–ECU change (b = +0.600), whereas Competence showed the smallest (b = −0.221) and decreased in 5 of 8 identities. GPT-4o showed the same overall direction of change across all five outcomes, and all five identity-level MS–ECU sign-flip tests remained significant after Holm correction. The predicted larger CU–ECU change was not supported in humans, and no GPT-4o outcome showed a significant transition difference. For Discomfort, the larger observed change occurred from MS to CU in humans but from CU to ECU in GPT-4o. Across both evaluators, tighter framing produced less favorable social impressions and greater Discomfort for the same neutral identities.

Hua Hwang, 강다현, Sung Park · 0 citations
#large language models Open access Sep 2026

Narrating Authority Through Storytelling in Entrepreneurial Podcast Interviews

How do entrepreneurs claim the authority to lead an enterprise that was never theirs alone? Although the discursive construction of leadership has attracted growing interest, most existing work is qualitative and case-limited, and leadership authority has rarely been examined as a discursive phenomenon at scale. Drawing on discursive, relational, and narrative leadership theory, this study analyzes 59 entrepreneurial podcast interviews from the Innovation Fuel series, recorded between 2020 and 2026, combining lexicon-based authority detection, sentence-embedding, topic modeling, and clustering. Under the stated lexical rules, soft classifications account for 89.8% of transcripts, and individually voiced and collectively voiced authority language routinely co-occur within the same narratives. Borrowed authority is used as a provisional interpretive label for legitimacy claimed individually yet drawn from teams and networks; it is offered as a testable interpretation rather than a separately validated construct. Recorded guest gender is neither associated with the episode-level collective ratio (p = .205) nor recoverable from the analyzed discourse out-of-sample (cross-validated AUC of about 0.59), and gender labels are interspersed in the embedding space. Because complete transcripts include host speech, these findings cannot be attributed to guest speech alone. Crucially, each apparent result is submitted to a matched null test: the soft-authority prevalence and the gender null survive, whereas apparent cluster and topic authority bands are shown to be aggregation artifacts over poorly separated groups. Distinguishing patterns that survive such tests from those that do not, the study models a transparent, self-auditing approach to computational discourse analysis and contributes a reproducible workflow for leadership communication research on large corpora of naturally occurring talk.

Gelareh Farhadien, Faezeh S. Aarabi, Dave Keighron · 0 citations
#large language models Open access Sep 2026

Explainable Hallucination Detection in Large Language Models Using Evidence-Grounded Verification

This paper presents Evidence-Grounded Verification (EGV), a modular framework for explainable hallucination detection in large language models. EGV decomposes model outputs into atomic claims, retrieves independent evidence, performs fine-grained four-way verification (Supported, Partially Supported, Contradicted, Unverifiable), and generates template-constrained explanations that are causally linked to the underlying evidence. Unlike most existing detectors that output only a scalar score or binary label, EGV treats explainability as a core architectural requirement. The framework is designed to run on consumer hardware using freely available models. A pilot study on 300 items (600 claims) from the HaluEval QA benchmark (using oracle evidence) compares three verification backends. Results show that a hybrid decision rule combining an off-the-shelf NLI model with simple lexical and length-based features substantially outperforms pure NLI (F1 rises from 3.5% to 65.6%), while a counterfactual evidence-swap test achieves a faithfulness score of 78.0%. The paper provides a full evaluation protocol, an honest analysis of limitations, and clear directions for future validation. This is a Methodology / Framework paper intended as a practical starting point for research on transparent, evidence-grounded hallucination detection.

Puwakpitiyage Sasmitha Mahesh Madhubhashana · 0 citations
#large language models Open access Sep 2026

Narrating Authority Through Storytelling in Entrepreneurial Podcast Interviews

How do entrepreneurs claim the authority to lead an enterprise that was never theirs alone? Although the discursive construction of leadership has attracted growing interest, most existing work is qualitative and case-limited, and leadership authority has rarely been examined as a discursive phenomenon at scale. Drawing on discursive, relational, and narrative leadership theory, this study analyzes 59 entrepreneurial podcast interviews from the Innovation Fuel series, recorded between 2020 and 2026, combining lexicon-based authority detection, sentence-embedding, topic modeling, and clustering. Under the stated lexical rules, soft classifications account for 89.8% of transcripts, and individually voiced and collectively voiced authority language routinely co-occur within the same narratives. Borrowed authority is used as a provisional interpretive label for legitimacy claimed individually yet drawn from teams and networks; it is offered as a testable interpretation rather than a separately validated construct. Recorded guest gender is neither associated with the episode-level collective ratio (p = .205) nor recoverable from the analyzed discourse out-of-sample (cross-validated AUC of about 0.59), and gender labels are interspersed in the embedding space. Because complete transcripts include host speech, these findings cannot be attributed to guest speech alone. Crucially, each apparent result is submitted to a matched null test: the soft-authority prevalence and the gender null survive, whereas apparent cluster and topic authority bands are shown to be aggregation artifacts over poorly separated groups. Distinguishing patterns that survive such tests from those that do not, the study models a transparent, self-auditing approach to computational discourse analysis and contributes a reproducible workflow for leadership communication research on large corpora of naturally occurring talk.

Gelareh Farhadien, Faezeh S. Aarabi, Dave Keighron · 0 citations
#large language models Open access Sep 2026

Explainable Hallucination Detection in Large Language Models Using Evidence-Grounded Verification

This paper presents Evidence-Grounded Verification (EGV), a modular framework for explainable hallucination detection in large language models. EGV decomposes model outputs into atomic claims, retrieves independent evidence, performs fine-grained four-way verification (Supported, Partially Supported, Contradicted, Unverifiable), and generates template-constrained explanations that are causally linked to the underlying evidence. Unlike most existing detectors that output only a scalar score or binary label, EGV treats explainability as a core architectural requirement. The framework is designed to run on consumer hardware using freely available models. A pilot study on 300 items (600 claims) from the HaluEval QA benchmark (using oracle evidence) compares three verification backends. Results show that a hybrid decision rule combining an off-the-shelf NLI model with simple lexical and length-based features substantially outperforms pure NLI (F1 rises from 3.5% to 65.6%), while a counterfactual evidence-swap test achieves a faithfulness score of 78.0%. The paper provides a full evaluation protocol, an honest analysis of limitations, and clear directions for future validation. This is a Methodology / Framework paper intended as a practical starting point for research on transparent, evidence-grounded hallucination detection.

Puwakpitiyage Sasmitha Mahesh Madhubhashana · 0 citations
#large language models Open access Sep 2026

Development and benchmark validation of PubChat for PubMed-grounded multilingual biomedical literature retrieval

The rapid growth of biomedical literature has rendered traditional systematic reviews unsustainable. Although large language models (LLMs) offer automation potential, citation fabrication and unreliable evidence discrimination remain critical barriers. Here, using PubChat as a PubMed E-utilities-grounded retrieval framework, we tested whether source-verifiable AI-assisted retrieval could preserve recall, criterion-driven relevance stratification, and multilingual accessibility in systematic-review benchmarking. The system comprises Phase I, hierarchical decomposition of the research question into five relevance levels, and Phase II, multi-round retrieval with embedding-based pre-filtering and three-round LLM verification. Validated against 20 Cochrane systematic reviews (585 ground-truth articles) across eight languages, PubChat was benchmarked against four general LLMs (GPT-5.2-Thinking, Gemini 3.0 Pro, Grok-4.1-Thinking, Qwen3-Max), one search-augmented retrieval tool (Perplexity-Sonar), and three specialized retrieval tools (Elicit, ASTA, and Consensus). PubChat produced no fabricated citations in this benchmark and showed distinct recall–precision profiles across its three operating modes. Among the evaluated configurations, PubChat-Broad achieved the highest observed recall and nDCG, whereas PubChat-Core achieved the highest observed F1- and F2-scores. In a user evaluation of 279 biomedical researchers across 18 countries, PubChat scored above 80/100 for reliability, innovation, efficiency, user experience, and self-reported preference over the evaluated alternatives (78% vs. specialized tools and 72% vs. manual search). An exploratory MIDE case study further illustrates how PubChat-derived corpora can be organized into evidence-traceable research-gap candidates for hypothesis prioritization. PubChat provides a benchmark-validated PubMed-grounded framework for source-faithful biomedical literature retrieval, relevance stratification, and structured evidence organization.

Xuenan Zhuang, Dan Cao, Ruoyu Chen et al. · 0 citations
#large language models Open access Sep 2026

Iterated Agent for Symbolic Regression

Abstract Symbolic regression (SR), the automated discovery of mathematical expressions from data, is a cornerstone of scientific inquiry. However, it is often hindered by the combinatorial explosion of the search space and a tendency to overfit. Popular methods, rooted in genetic programming, explore this space syntactically, often yielding overly complex, uninterpretable models. This paper introduces IdeaSearchFitter, a framework that employs Large Language Models (LLMs) as semantic operators within an evolutionary search. By generating candidate expressions guided by natural-language rationales, our method biases discovery towards models that are not only accurate but also conceptually coherent and interpretable. We demonstrate IdeaSearchFitter's efficacy across diverse challenges: it achieves competitive, noise-robust performance on the Feynman Symbolic Regression Database (FSReD), outperforming several strong baselines; discovers mechanistically aligned models with good accuracy-complexity trade-offs on real-world data; and derives compact, physically-motivated parametrizations for Parton Distribution Functions in a frontier high-energy physics application. IdeaSearchFitter is a specialized module within our broader iterated agent framework, IdeaSearch, which is publicly available at \href{https://www.ideasearch.cn/}{https://www.ideasearch.cn/}.

Zhida Song, Zeyu Cai, Shutao Zhang et al. · 0 citations
#large language models Open access Sep 2026

Fixing, Breaking, or Faking It? An Execution-Calibrated Evaluation of LLM Vulnerability Patching in JavaScript, Python, Go, and Java, and the Limits of LLM-as-Judge

Large language models (LLMs) increasingly repair software vulnerabilities, but most evaluations judge only similarity to a developer fix or removal of the weakness. Neither reveals whether working code was broken. We evaluate eight commercial and open LLMs on 922 JavaScript vulnerability patches, scoring neutralisation and functional preservation. Lacking tests, we score at scale with a reference-based LLM judge, calibrated against execution on a 144-patch benchmark and 254 Java-CVE patches, plus a cross-family judge. The best model fixes 23% of vulnerabilities (judge-based), and cost-efficiency inverts the accuracy ranking. Our central finding concerns the instrument: both judges flag more over-fixes than execution confirms (precision 5–10%), yet on the functional axis agree far more with each other (κ=0.75) than with execution (κ≤0.26), so judge–judge agreement measures reliability, not validity. On real Java code the over-reporting persists, while the judges’ correctness estimates diverge, leaving no single judge trustworthy. Over-fixing is genuine but, under execution, uncommon: a few percent of vulnerability-removing patches, or under 2%, excluding one artefact-prone scenario, both lower bounds. Only adequately tested execution measures the functional-preservation rate, so security-patch evaluation must run the code, use a judge only to rank models, and weigh costs. We release the harness and executable benchmark.

Patrick Deininger, Wolfgang Slany · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.