Skip to content

Category

large language models

450 papers

#large language models Open access Sep 2026

Phase-Based Dynamic Prediction of Delayed Cerebral Ischemia After Aneurysmal Subarachnoid Hemorrhage: Comparison of a Large Language Model with Intensive Care Specialists

Background and Objectives: Delayed cerebral ischemia (DCI) is a major determinant of poor outcome after aneurysmal subarachnoid hemorrhage (aSAH), yet early risk prediction remains difficult, particularly in sedated or ventilated intensive care patients. We evaluated whether a large language model (LLM) could predict DCI from phase-based dynamic clinical data, compared with intensive care specialists. Materials and Methods: In this single-center, retrospective study, 216 consecutive patients with aSAH were assessed at three predefined phases of accumulating clinical data (day 1; days 1 + 3; days 1 + 3 + 5). For each patient–phase, an LLM (ChatGPT, GPT-5.5 Thinking) and two blinded intensive care specialists predicted DCI risk using only the data available up to that time point. DCI was adjudicated by a blinded three-member panel. Discrimination was assessed by the area under the receiver operating characteristic curve (AUC), with non-inferiority defined a priori as Δ = 0.10. Calibration, decision-curve analysis, and reproducibility were also assessed. Results: Of 216 patients, 60 (27.8%) were DCI-positive and 15 were indeterminate; the primary sample comprised 201 patients. In the prespecified primary analysis, the LLM met the non-inferiority criterion relative to both specialists across all three phases (LLM AUC 0.703–0.747; specialists 0.712–0.764); however, in an equal-granularity sensitivity analysis, non-inferiority remained supported only in Phases 2 and 3 and was not demonstrated in Phase 1. Discrimination increased numerically as data accumulated (Phase 1 vs. 3, p = 0.072), an increase that was attenuated in a landmark-restricted analysis accounting for DCI-onset timing. The LLM showed the lowest false-reassurance rate (3.3–11.7%), reflecting a more cautious threshold rather than better discrimination. Confidence did not reliably indicate accuracy (~25% of high-confidence predictions were wrong); outputs were highly reproducible (Fleiss κ 0.887; intraclass correlation coefficient (ICC) 0.968). Conclusions: An LLM achieved DCI discrimination that was non-inferior to—but not better than—that of experienced specialists; its unreliable confidence scores support clinician-supervised rather than autonomous use.

Mustafa Ay, Tüfek Öztan Dilara, Şule Asri et al. · 0 citations
#large language models Open access Sep 2026

Optimization of Machine Learning–Based Recommendation Systems on E-Commerce Platforms

This study aims to analyze optimization strategies for machine learning–based recommendation systems in e-commerce environments, identify commonly applied algorithms, and examine emerging opportunities and implementation challenges. A Systematic Literature Review (SLR) was conducted following the PRISMA 2020 framework. Literature was collected from six major academic databases, covering publications from 2020 to 2025. From an initial pool of 286 records, 10 studies met the eligibility criteria and were included in the final analysis. The findings indicate that deep learning, hybrid recommendation models, sequential recommendation approaches, and large language model–based systems significantly enhance recommendation accuracy and personalization. User behavior analytics emerged as a critical factor in adaptive recommendation systems, while conversational AI and multimodal technologies represent promising future directions. Despite these advancements, issues related to scalability, explainability, fairness, and privacy remain significant challenges requiring further research and optimization.

Stephen Gregorius Kurnia, Muhammad Rizki Perdana, Aldian Yusup · 0 citations
#large language models Open access Sep 2026

Idiomatic Creativity and Emergent Fixedness: Gradient Lexical Specification in Large-Scale Cross-Linguistic Corpora

Abstract This paper demonstrates that idiomatic fixedness is manipulation-dependent: the constructional elements that resist omission in idiom variants are systematically different from those that resist replacement. Using a manually verified dataset of over 14,000 idiomatic variants across 25 idioms in three languages (English, Japanese, and Korean), drawn from web-based TenTen corpora, we compare two types of structural manipulation – contraction, in which a canonical element is omitted, and substitution, in which it is replaced – to examine how figurative interpretation is preserved under each type of pressure. Results show that the set of retained constructional elements differs systematically between contraction and substitution, indicating that, for most longer idioms, no invariant core remains stable across manipulation types. Patterns that appear categorically fixed in intuition-based accounts instead emerge as context-sensitive stability effects visible only through large-scale digital aggregation of naturalistic data. For instance, in substitution variants like let the agenda out of the bag (canonical: let the cat out of the bag ), the canonically central element cat is sacrificed to ground reference in discourse, despite agenda ’s low collocational potential with bag . Analyzing retention asymmetries, we propose a refinement of the constructional continuum. We introduce a model where schematicity and lexical specification vary as partially independent dimensions, providing a framework that accounts for the negotiability trade-off observed when speakers balance figurative recognition with discourse-specific needs. Within this framework, fixedness is reconceptualized as context-dependent stability – an emergent equilibrium shaped by constructional and discourse constraints and usage frequencies rather than an inherent property of particular lexical items. Corpus methods thus provide empirical leverage for refining the theoretical architecture of Construction Grammar.

Carey Benom, Youngmin Oh · 0 citations

ChatSeven: An Agentic AI-Based Multi-Agent Platform for Multi-Channel Customer Conversation Management and Campaign Automation

Abstract—Customer engagement platforms increasingly re-quire artificial intelligence (AI), multi-channel messaging, work-flow automation, and outbound campaign delivery within a single operational system. Traditional conversational agents—including rule-based chatbots, retrieval-augmented generation (RAG) only bots, and standalone large language model (LLM) interfaces—typically address only a subset of these requirements. This paper presents ChatSeven, an agentic AI-based multi-agent platform for multi-channel customer conversation management and campaign automation, implemented as the ChatSeven production system. ChatSeven integrates LangGraph ReAct agents, PGVector-based RAG, a unified inbox across Web Chat, WhatsApp, Email, SMS, and Instagram, visual workflow Flows, and queue-based campaign broadcasting under a master–regional multi-tenant database architecture. The platform is realized through a React frontend, Express/TypeScript backend, FastAPI Chat and Vector microservices, PostgreSQL, Redis/Bull job queues, and Socket.IO real-time messaging. Experimental evaluation on AI answer quality, RAG retrieval, tool reliability, flow completion, response latency, and campaign delivery demonstrates that ChatSeven closes a critical integration gap between agentic LLM research and enterprise conversation operations. Demo evaluation reports approximately 88% AI answer accuracy, 0.86 Precision@5 for RAG retrieval, 2.4 s average response latency, 90–94% tool success, 84% flow completion, and 96% campaign delivery. Index Terms—Agentic AI; Multi-agent systems; Conversa-tional AI; Retrieval-Augmented Generation; Omnichannel cus-tomer engagement; Campaign automation; LangGraph; Multi-tenant SaaS.

Afnan Shaikh, Jeetendra Singh Yadav · 0 citations
#large language models Open access Sep 2026

Exploring the Development of AI-Mediated Competence for Sustainable Translation and Interpreting Education: A Bibliometric Review

Artificial intelligence has become increasingly central to translation and interpreting education, shaping classroom practice, feedback, assessment, and professional preparation. Yet the literature remains fragmented across work on machine translation, computer-assisted translation, post-editing, automatic speech recognition, generative AI, large language models, and AI-supported assessment. This bibliometric review examines how research in this area has developed from 2014 to May 2026, with particular attention to the movement from tool-oriented technology training toward AI-mediated competence. Bibliographic records were retrieved from Web of Science and Scopus and analysed using VOSviewer and CiteSpace. The analysis focuses on publication trends, collaboration patterns, keyword co-occurrence, thematic clusters, and keyword bursts. The results show limited output before 2018, steady growth between 2019 and 2021, and rapid expansion after 2022. The keyword evidence points to a shift from machine translation, post-editing, and translation technology training toward AI literacy, evaluative judgement, output verification, feedback practices, ethical responsibility, professional agency, and human-AI collaboration. Interpreting-related research is still less developed than translation-oriented research. The findings suggest that AI integration in translation and interpreting education is not simply a matter of adopting new tools, but part of a broader process of competence development.

Qing Zheng, Mansour Amini · 0 citations
#large language models Open access Sep 2026

Reimagining biomedical science workflows in the age of large language models

Abstract Large language models (LLMs) are generative artificial intelligence (AI) models that are rapidly reshaping the practice of biomedical science. Their ability to synthesize literature, generate analytical code, and interface with multimodal data offers a new framework for accelerating discovery. Yet their integration into scientific workflows remains irregular, and the field lacks clear guidance for reliable and productive use. We review emerging evidence on researcher adoption, highlight common failure modes such as so-called hallucinations (i.e. confabulations) and overgeneralization, and provide practical recommendations for domain-informed use of LLMs in basic biomedical research. We structure this review around four domains in which LLMs increasingly augment biomedical science: administrative tasks, literature search and synthesis, data analysis, and scientific writing. For each domain we provide practical guidance, illustrative use cases, and examples of free or low-cost tools that researchers can readily adopt. Finally, we discuss the organizational and cultural changes for biomedical science to leverage LLMs responsibly, including transparent reporting, human-in-the-loop validation, and alignment with scientific rigor and reproducibility standards. Together, these recommendations provide a path for integrating LLMs into biomedical research in ways that enhance, rather than replace, human expertise and accelerate the path from biological insights to beneficial human impact.

Kaleigh F. Roberts, Srinivas Koutarapu, Justin Melendez et al. · 0 citations

Mechanistic understanding and peptide ingredient screening for type III collagen via biological knowledge graph and molecular docking

Abstract Objective Type III collagen plays a key role in maintaining skin elasticity and dermal resilience, yet its decline is associated with visible signs of skin ageing. This study aimed to establish a data‐driven approach to identify cosmetic peptides with potential to enhance type III collagen through multi‐target regulation. Methods A combined strategy integrating natural language processing (NLP) and knowledge graph (KG) analysis was applied to systematically identify targets associated with type III collagen regulation. Peptides listed in the Inventory of Existing Cosmetic Ingredients in China (IECIC) were screened based on structure‐based molecular docking with multiple targets. Candidate peptides were selected for experimental evaluation using human dermal fibroblasts and ex vivo skin models. Results By integrating large‐scale literature mining with NLP and KG‐based gene interaction analysis, we constructed a comprehensive network of genes involved in type III collagen regulation, covering processes such as collagen synthesis, degradation and extracellular matrix (ECM) organization. Based on this network, candidate peptides with predicted multi‐target interactions were prioritized through structure‐based screening. In vitro, four selected peptides—palmitoyl tetrapeptide‐7, acetyl hexapeptide‐8, palmitoyl tripeptide‐1 and palmitoyl tetrapeptide‐10—significantly increased type III collagen levels following UV exposure, without detectable cytotoxicity. The combination of these peptides exhibited a greater effect than individual components. In ex vivo skin tissues, the peptide combination further enhanced the levels of type I, III and IV collagens, indicating a broader modulatory effect on dermal ECM components. Conclusion This study provides a practical framework for the identification of cosmetic peptides targeting collagen regulation. The findings support a multi‐target approach and suggest that peptide combinations may enhance collagen‐related outcomes relevant to skin ageing applications.

S. L. Zang, Zhiqiang Chen, Jie Qiu et al. · 0 citations

Comparing Human‐Created and NotebookLM‐Generated Podcasts for EAP Instruction: Differences in Linguistic Complexity and the Role of Prompt Design

ABSTRACT Science popularization podcasts are valuable resources for EAP instruction because they make academic content more accessible to learners. However, producing high‐quality podcasts is time‐consuming and resource‐intensive, and existing resources may not always align with instructional goals or learners’ needs. NotebookLM, a large language model (LLM)‐based tool, may offer possibilities for generating multimedia materials more efficiently. Yet, little is known about how linguistically comparable such AI‐generated podcasts are to human‐created ones or how prompt design may shape their linguistic complexity for pedagogical use. To address this gap, this study examined the lexical and syntactic complexity of human‐created and NotebookLM‐generated podcasts derived from the same research‐article source texts. Four corpora were analyzed: a corpus of human‐created podcasts from the Nature Podcast series and three corpora of NotebookLM‐generated podcasts produced under different conditions: the default setting (no prompt) and prompts intended to approximate CEFR B1‐ and B2‐level output for EAP instruction. Fourteen indices of lexical and syntactic complexity were analyzed. The findings showed that NotebookLM‐generated podcasts were generally more lexically complex but less syntactically complex than human‐created podcasts. Prompt design also differentiated the linguistic complexity of NotebookLM‐generated output: the B1‐level prompt generally produced less complex output than the B2‐level prompt, while the default setting generated the most challenging materials. Although the prompted outputs were not fully comparable to the human‐created materials, they showed greater similarity on several linguistic features. These findings provide preliminary evidence of NotebookLM's potential for EAP materials development and highlight the importance of prompt design in shaping AI‐generated pedagogical materials.

Chen-Yu Liu · 0 citations

DEL-GPT: learning the language of DNA-encoded libraries to design focused screening collections

Focused DNA-encoded libraries (DELs) yield higher hit rates and cleaner selection data than billioncompound collections, and are affordable enough to build for a single target. Designing one means choosing a few hundred synthons from a much larger pool, and in the split-and-pool format that choice is collective and irreversible: the library is the full combinatorial product, so each synthon's value depends on every other one chosen alongside it. Brute force is out of reach – there are on the order of 10189 ways to draw 100 synthons from 3,000. We therefore recast the problem as sequence modeling. Treating a synthon as a token and an efficiency-ordered synthon set as a sentence, a generative pretrained transformer trained by next-token prediction learns context-dependent synthon value and composes new sets one synthon at a time. DEL-GPT was trained on one million ranked synthon sets drawn from NADEL, a validated 58,302-member library, with one model per target for five NAD+-dependent enzymes (PARP1, PARP2, PARP10, PARP12, PARP15). Across all five, DEL-GPT libraries outperformed those built from randomly drawn synthons – a demanding baseline, since every NADEL synthon is an expert-selected NAD+ mimetic – and contained two to three times more compounds that experimental affinity selection retained. Quality falls off once generated sets exceed the size range seen in training, but a model trained on longer sequences generated correspondingly larger high-quality sets – a validated recipe for scaling. The method is chemistryagnostic and applies to any combinatorial library over a shared building-block set.

Akhila Mettu, Naveed Naemi, Raphael Franzini et al. · 0 citations
#large language models Open access Sep 2026

A review of bias detection and fairness auditing techniques in LLMs

The rise of large language models (LLMs) has sparked worries about inherent social biases and issues related to fairness. Earlier studies have investigated bias identification in word embeddings, interventions aimed at fairness in algorithms, and frameworks for auditing at the system level. Nonetheless, these methods remain disorganized, with variations in datasets, evaluation methods, and implementation processes. In this paper, we provide a thorough literature review to encapsulate prior research on bias identification and fairness auditing, categorizing the findings according to various stages of study. Additionally, we analyze the limitations in coverage and consistency of widely used benchmark datasets. To tackle these issues, we propose a unified pipeline for dataset integration and a modular framework for bias auditing. Recognized significant research gaps include the absence of intersectional bias modeling, a shortage of standardized evaluation metrics, and challenges in scalability for real-time auditing systems.

Nani Kartik Kaveti, Tanuja Pattanshetti · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.