Background & Aims: Multi-target stool DNA (mt-sDNA) is an established stool-based colorectal cancer screening test, yet long-term performance is not well-described. We evaluated follow-up colonoscopy rates & quality, neoplasia yield, and repeat-testing adherence over 10 years with an integrated screening program. Methods: We assembled a retrospective cohort of mt-sDNA tests (September 2014–July 2025) within a multi-site health system. A validated large-language-model pipeline extracted colonoscopy, pathology, and quality endpoints from free-text reports. Detection outcomes (adenoma detection rate, advanced neoplasia, modified advanced adenoma, sessile serrated lesion (SSL), and negative exams) were calculated among high-quality exams; adenocarcinoma, bowel preparation adequacy, and cecal intubation rate (CIR) used all screening exams. We assessed colonoscopy completion ≤1 year after positive mt-sDNA, follow-up colonoscopy findings, findings by time-to-colonoscopy, and adherence to repeat mt-sDNA after negative result. Results: Among 148,624 tests in 110,410 individuals, positivity was 13.6%. Of 19,506 positives, 65.1% completed colonoscopy within 12 months (median 82 days). Colonoscopy after a positive mt-sDNA had high yield: advanced adenoma diagnosis was 30.7%; adenocarcinoma was 1.3%. Of patients who had colonoscopy within 3 years of a negative mt-sDNA test, 45.2% had adenoma/serrated lesion, 9.3% had advanced adenoma and 0.5% had colon cancer. Adenoma detection rate after positive mt-sDNA (>50%) exceeded specialty society quality benchmarks and increased over time; Longer intervals (>6 months) to colonoscopy after a positive test were associated with higher advanced adenoma (trend P=0.001).After an initial negative mt-sDNA, only 35.3% were screened again by mt-sDNA testing within 3 years. Conclusions: Over a decade,mt-sDNA screening achieved high colonoscopy qualitybut revealed critical gaps in the care cascade. One-third of positive tests lacked timely follow-up and nearly two-thirds of negative tests were not repeated on schedule, suggesting opportunities for system-level interventions.
Sushil Kumar Garg, Carley F. Hintz, Hannah J. Kolarik et al.· The American Journal of Gast...· 0 citations
This article examines the impact of the digital environment on the status and place of language. The focus is on the concept of digital linguistic sovereignty, understood as a nation’s ability to ensure the functional competitiveness of the language in the age of artificial intelligence and big data. In this regard, Kazakhstan presents a unique case study of a country transitioning from traditional language protection to the active construction of digital sovereignty. As a state undergoing active nation-building, the country simultaneously demonstrates one of the highest levels of digitalisation in the post-Soviet space. The implementation of large-scale state programmes (such as “Digital Kazakhstan”) has led to the formation of a new communicative environment – a virtual linguistic landscape – in which the dynamics of interaction between the state language and global languages are assuming qualitatively different forms. The methodological basis of the study is an interdisciplinary approach combining critical discourse analysis of official documents, institutional analysis of technological projects such as the National Corpus of the Kazakh Language, the Coursera platform in Kazakh, the translation of educational materials for higher education institutions, and others, as well as the study of expert discourses in new media. The study’s findings illustrate how the state regulates the process of language planning in the context of increasingly pervasive technological advancement, transitioning from a protective model toward digital modernisation. The creation of mass digital content in academic and technical fields has been demonstrated to contribute to the growing intellectual recognition of the Kazakh language. The authors of the study argue for the need to transition from formal status planning to an inclusive model of digital citizenship in order to achieve “linguistic justice.” It has been established that language digitalization is becoming a tool of “soft power,” transforming civic identity and expanding access to digital capital. The practical significance of this study lies in the potential application of the proposed analytical framework to the study of language processes in other multilingual societies undergoing digital transformation.
Sholpan Zharkynbekova, Gulbagira Ayupova, Bakhyt Galiyeva et al.· Frontiers in Psychology· 0 citations
Executive Summary This work is a literature review on the task of closed-domain event extraction, a task within natural language processing (NLP) where the goal is to detect the presence of an event (an occurrence of an action or state change) and event-related information (known as arguments) and to return the event structure. In our review, we focus on recent state-of-the-art (SOTA) approaches which can be applied to unstructured English-language text. After the introduction, we begin by providing an overview of the terminology of closed-domain event extraction, including a breakdown of the constituent sub-tasks. Next, we discuss the dataset requirements for this task, and note the core datasets used in the literature. We begin our review of the literature by providing a typology of event extraction approaches based on the common differences between extant approaches. We provide a table summarising the main results in the literature and discuss standards of evaluation in the literature. We found that all reviewed SOTA approaches utilise a PLM, with the choice of model often dependent on whether a classification or generative approach is taken-no single approach is dominant. BERT and BART are popular choices of PLM in the literature. We argue that evaluation standards have been inconsistent, resulting in results which are difficult to compare. We then describe some of the details of the most performant models in our table. Finally, we discuss some of the main themes we have identified in our review. In particular, we outline some clear issues in the extant literature. We argue that there is a need for a new, open source dataset to act as the primary benchmark to facilitate more academic research, with high inter-annotator agreement (IAA) ensuring that good performance on the benchmark is meaningful. We also call for standardisation around pre-processing and evaluation of results to ensure the comparability of results. We note that low-resource performance has typically been under-explored. In our recommendations for future work, we argue that there is clear potential to assess the performance of modern, causal large language models (LLMs) such as the Llama or GPT model families. We argue that it there is clear potential to explore both in-context learning and fine-tuned approaches with these models. We also argue that these models have clear potential applications in the generation of synthetic data. We conclude by summarising our main points. 1
Joanna Cameron Knight, Phil Swatton, Alex Hickey et al.· Alan Turing Institute Resear...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
The goal of question generation is to automatically produce relevant and meaningful questions from diverse inputs such as knowledge bases, natural language texts, and images. With the rapid advancement of neural architectures, neural question generation (NQG) has attracted growing attention across both academia and industry. In this survey, we provide a comprehensive review of developments in NQG, spanning traditional neural approaches to the latest paradigms driven by large language models (LLMs) and multimodal large language models (MLLMs). We begin by outlining the fundamental components of NQG, including its problem formulation, benchmark datasets, evaluation metrics, and representative applications. Next, we categorize existing methods into three main types: structured NQG , which relies on structured data sources; unstructured NQG , which handles loosely structured inputs such as texts or images; and hybrid NQG , which integrates multiple modalities. For each category, we review representative neural models and synthesize the problems addressed by successive generations of methods, their remaining limitations, and the motivations behind major methodological transitions. Furthermore, we trace the progression of NQG from supervised neural approaches and pre-trained models to prompting, retrieval-augmented generation, reinforcement learning, and emerging tool-augmented and agent-based paradigms. We also discuss how recent LLMs and MLLMs have enabled more contextually aligned, knowledge-grounded, and reasoning-enhanced question generation, together with emerging concerns such as hallucination, bias, and evaluation reliability. Finally, we outline open challenges and emerging research trends, offering a forward-looking perspective on the evolution of NQG. This survey presents a meticulously curated compilation of related papers, datasets, and code, serving as a comprehensive resource for anyone studying NQG.
Large Language Models (LLMs) have demonstrated remarkable general-purpose abilities across a wide range of domains, and these strengths have also been increasingly evidenced in recommender systems. However, existing methods that attempt to integrate collaborative signals into LLMs often fail to preserve their foundational knowledge. This loss is critical in text-rich recommendation, where robust semantic understanding is required to interpret user reviews and item profiles. We propose PALRec , a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed. We first construct evidence-grounded user and item profiles from reviews and use them as concise pseudo-labels for reconstruction. We then introduce lightweight, trainable user and item embedding modules optimized with a multi-task objective that combines next-item prediction and profile reconstruction. These modules are trained jointly to align collaborative signals with the LLM’s semantic space without modifying the backbone. We also employ token-aware loss decomposition and frequency-aware reweighting to stabilize training and mitigate popularity bias. Experiments on public benchmarks show that PALRec consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge. This result highlights that maintaining the LLM’s semantic understanding is crucial for effectively exploiting textual information in recommender systems.
Hyunsoo Na, Minseok Gang, Sang‐goo Lee et al.· ACM Transactions on Informat...· 0 citations
Understanding protein mechanisms in health and disease requires characterizing the functional roles of individual amino acid residues. To explore the role of residues and their mutations, we have developed Atlantis, a database that integrates structural and functional information at the human proteome residue level. A graph database enables complex queries and the retrieval of integrated information for multiple functional analysis of protein systems. A Model Context Protocol (MCP) connector allows the interrogation of the resource through Large Language Models (LLMs) or agentic frameworks for biomedical research. Atlantis annotates over 11M residues across 20k human proteins, identifying hundreds thousands intra- and inter-protein contacts in PDB as well as AlphaFoldDB structures. We also provide the possibility to analyze and integrate predicted 3D complexes inputted by the user, and we showcased these features on hundreds of AlphaFold-multimer complexes of GPCRs and LRRK2 interaction networks. The tool is freely accessible at https://atlantis.bioinfolab.sns.it/.
Natalia De Oliveira Rosa, Piergiorgio Ferronato, Martina Varisco et al.· bioRxiv (Cold Spring Harbor...· 0 citations
The expansion of smart technologies has transformed teaching and learning paradigms in physical education and sport sciences, paving the way for interactive, data-driven, and personalized instruction. Objective: This study aimed to examine the role of smart technologies in transforming physical education instruction and to identify the associated opportunities and challenges. Methodology: A narrative–analytical review approach was adopted, examining studies published between 2020 and 2026 in the domains of artificial intelligence (AI), wearable technologies, the Internet of Things (IoT), virtual and augmented reality (VR/AR), gamification, and large language models (LLMs). Findings / Results: The findings revealed that smart technologies enhance learning quality, motivation, and athletic performance by enabling personalized learning, instant feedback, intelligent assessment, and performance data analytics. Nevertheless, data privacy and security, high implementation costs, the digital divide, algorithmic accuracy, and teacher readiness represent the most prominent challenges. Conclusion & Implications: Smart education achieves optimal effectiveness only when integrated with teacher competence, appropriate technological infrastructure, and ethical considerations. The purposeful implementation of these technologies can foster a dynamic, personalized, and data-driven ecosystem for physical education.
Reza Zolghadri, Amirmohammad Naderkordi· Zenodo (CERN European Organi...· 0 citations
Companion code and data for "Multi-vendor evaluation of large language models for ACMG/AMP variant classification with controlled data contamination" (Journal of Genetics and Genomics). 9 LLMs × 5,000 temporally-blinded ClinVar variants (45,000 evaluations), with sub-experiments on allele-frequency ablation, conflicting interpretations, MaveDB functional evidence, determinism, and prompt symmetry. Changes since v1.0.0: fourteen rounds of integrity audit — all reported numbers recomputed from raw data, figures regenerated data-driven, references verified field-by-field against Crossref/arXiv, repository now self-contained (figures reproduce byte-identically from archived data); automated verification suite included (final_gate.py, table_cell_audit.py, ref_full_audit.py).
zksdu· Zenodo (CERN European Organi...· 0 citations
This chapter traces the emergence of a new computational paradigm shaped by advances in neural networks, generative models, and multimodal intelligence. It outlines the evolution from expert systems – rule-based and deterministic – to adaptive learning systems capable of inference and creativity. The rise of architectures such as transformers and large language models redefines computation from retrieval to generation, enabling machines to reason across textual, visual, and spatial modalities. In this framework, computation evolves from a closed process of execution to an open, generative ecology capable of modelling, simulating, and reasoning about the world. The chapter situates this shift as both technological and epistemological, redefining how creativity, problem-solving, and design knowledge are constructed in the age of intelligent computation.
With the advancement of Generative Artificial Intelligence (GenAI) and in particular Large Language Models (LLMs), increasing focus has been placed on the development of conversational agents across domains, including education. Within education research, pedagogical agents (PAs) have traditionally been designed by researchers, producing positive results when implemented in ways consistent with good pedagogical practice. Findings suggest that the effectiveness of PAs depends on the degree to which learners relate to their agents and perceive them as socially and behaviorally realistic. One way to accomplish this is to include end users in the PA design process. Authoring tools have lowered technical barriers, enabling novice designers to participate in educational technology design. Building upon this work, we introduce the Pedagogical Agent Toolkit (PATK), an extendable and modular toolkit with AI-based features with the purpose of expanding novice designer agency and opportunity in the creation and customization of PAs. This paper outlines the design principles, architecture, and features of the PATK.
Samuel Hum, Jennie Lee, Jessica R. Gladstone et al.· 0 citations
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.