The rapid expansion of the oncology literature has outpaced manual curation of clinically relevant gene-cancer-drug associations and oncogenic driver evidence. Existing automated approaches often lack transparency or are difficult to scale across heterogeneous data sources. To address this gap, we developed megaMine, a transparent, rule-based, and context-aware literature-mining framework that integrates therapeutic and driver evidence from PubMed, PubTator, and Europe PMC by combining entity recognition, hierarchical heuristics, and contextual labeling. In therapy mode, megaMine was applied to approximately 100,000 oncology articles published between 2015 and 2025, yielding more than 23,000 structured sentence-level evidence records, with standardized annotations for drug response, resistance, and study context. Internal evaluation of context labels showed the strong separability between efficacy and non-efficacy evidence using ridge logistic regression (AUROC = 0.915; AUPRC = 0.941). Benchmarking against NCI/OncoKB-supported drug-cancer associations showed that curated clinical associations had higher megaMine composite evidence scores than unlabeled comparison pairs [median (IQR): 25.6 (9.07-72.5) vs. 3.61 (1.69-8.69); Wilcoxon rank-sum test, P < 2.2 × 10−16]. In driver mode, megaMine retrieved mutation- and biomarker-related evidence from an ERBB-focused gastric cancer query, generating 750 evidence rows from 200 PMIDs. These results demonstrate that deterministic and interpretable approaches can support scalable evidence extraction for downstream applications such as knowledge graph construction and literature-based evidence synthesis.
The increasing use of tumor sequencing has intensified the need for fast, traceable interpretation of genomic variants. General-purpose large language models can produce fluent answers, but unsupported statements, weak provenance, and stale knowledge limit their suitability for clinical genomics. We developed OncoGenRAG, a research framework that combines a parameter-efficiently fine-tuned BioBERT classifier with an entity-aware retrieval system over a curated, multi-source oncology knowledge base. The reported knowledge base contains 933 harmonized records derived from CIViC, ClinVar/dbSNP, Open Targets, UniProtKB/Swiss-Prot, Ensembl Variation, and linked PubMed literature. The classifier assigns one of five labels: Pathogenic, Likely Pathogenic, Variant of Uncertain Significance, Benign, or Oncogenic; the retrieval component ranks evidence records using subword TF-IDF similarity and explicit gene, variant, and cancer-type matches. A rejection rule suppresses answers when retrieval support is below a prespecified threshold. In the authors’ held-out evaluation, the classifier achieved 92.40% accuracy, 93.15% weighted precision, 92.40% weighted recall, and 92.65% weighted F1 score. In a separate benchmark of 100 clinical-style queries, OncoGenRAG achieved reported Precision@1 of 94.5%, Precision@3 of 96.8%, and 100% database grounding. No hallucinated answer was observed under the study’s operational definition, compared with a 41.0% no-hallucination rate for the ungrounded baseline. These results should be interpreted as internal validation rather than proof of universal safety because query construction, annotator agreement, class-specific performance, calibration, and external validation data were not available for independent analysis. OncoGenRAG provides a transparent design for evidence retrieval and abstention, but it is a research prototype and must not be used to select treatment without expert review.
Amaan Arif, José Valentim dos Santos Filho· bioRxiv· 0 citations
Background: Precision oncology relies on accurate interpretation of tumour-detected gene variants, to guide personalized treatment decisions. However, accurate interpretation of variants in context requires extensive information that is often buried within unstructured biomedical literature and obscured by inconsistent nomenclature, making manual retrieval labour-intensive and prone to omissions. Methods: To address this challenge, we developed Variantscape, a large-scale, automated pipeline and open-access web tool. It integrates traditional natural language processing methods with state-of-the-art large language models to extract, standardize, and analyze co-associations between genetic variants, cancer types, and therapeutic interventions from published biomedical abstracts. Findings: From over 3 million abstracts screened, 335,817 gene name-containing articles were eligible for downstream extraction. Among these, 7,423 (2.2%) simultaneously mentioned a variant, cancer type, and therapeutic agent, encompassing 3,902 unique variants across 98 cancer types and 388 therapeutic agents. This highlights the inefficiency of manual literature retrieval in molecular tumour board (MTB) workflows. Network analysis revealed 14,831 statistically significant co-associations, represented in a literature-derived graph with 4,388 nodes and 46,943 edges. Canonical alterations in well-studied cancers (e.g., BRAF V600E in melanoma) were strongly linked to established treatments, while several rare variants also emerged with high-confidence literature support. Interpretation: By applying large language models to biomedical literature, Variantscape enables scalable, context-aware extraction of trilateral variant-treatment-cancer relationships. This approach supports early evidence synthesis/hypothesis generation, highlights underrecognized or rare associations, and offers a practical resource for accelerating discovery and supporting precision oncology research and translation. Unlike static databases, Variantscape is continuously updatable and leverages large language model-based inference to uncover putative associations without manual curation. Variantscape has the potential to support MTB workflows and translational research by rapidly revealing signals from underlying abstracts.
M. Wosny, A. Blindu, M. Boesch et al.· medRxiv· 0 citations
Transcriptome-wide association studies (TWAS) can identify genes where genetically predicted gene expression is associated with disease risk, but translating those signals into therapeutic opportunities remains time-consuming, manual, and difficult to reproduce. We developed TRACE (TWAS-driven Repurposing through AI-assisted Curation of Evidence), a gene- and phenotype-agnostic computational pipeline that accepts a TWAS gene and effect-size direction, normalizes the gene symbol, retrieves FDA-approved drug-gene candidates from four online resources, collects related peer-reviewed literature from PubMed, and uses a fine-tuned biomedical language model to classify whether the literature supports a direct drug-gene relationship, the mechanism of action, and the direction of effect. The pipeline then compares the drug-derived direction with the direction implied by the TWAS effect estimate to rank candidate therapeutic pairs and flag potential drug safety concerns. The local classifier, built on BiomedBERT, was trained using pipeline-derived labels, BioCreative VI ChemProt gold-standard chemical-protein relation examples, and author-reviewed active-learning cases, reaching a held-out macro F1 of 0.809 across three simultaneous classification tasks. We validated the pipeline against a manually curated endometriosis gold standard of 43 drug-gene pairs spanning six TWAS-identified genes, developed through S-PrediXcan analysis of endometriosis GWAS summary statistics, manual querying of four drug-gene interaction databases for each gene, literature review of drug-gene mechanistic evidence, and Mendelian randomization validation of candidate pairs. External validation used two independently published genetically informed drug-repurposing studies in metabolic dysfunction-associated steatotic liver disease (MASLD) and type 2 diabetes (T2D). The pipeline recovered 90.7% of endometriosis pairs, 88.2% of MASLD pairs, and 92.9% of T2D pairs that were present in at least one queried database. Applied to 99 endometriosis-associated TWAS genes, the pipeline identified 1,089 FDA-approved drug-gene pairs, 32 candidate therapeutic pairs, and 77 potential safety concerns, including independent recovery of leuprolide acetate, an established endometriosis therapy. This framework provides a scalable, literature-grounded bridge from TWAS discovery to prioritized therapeutic hypotheses, while preserving uncertainty through manual-review flags and requiring downstream Mendelian randomization, electronic health record-based validation, and experimental follow-up before clinical interpretation.
C. O. Otieno, H. Seagle, A. Akerele et al.· medRxiv· 0 citations
Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying and confirming novel microbial oncogenicity could yield strategies and tools that will reduce disease burdens. However, relevant evidence may be dispersed across a vast biomedical literature that is infeasible for humans to comprehensively synthesize. Large Language Models (LLMs) may enable scalable, expert-level systematic evidence synthesis to identify high priority microbe-cancer pairs; however, such capabilities have not yet been demonstrated.
Domain experts were recruited to create a human-validated test dataset to benchmark the performance of LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, and GPT-5 Nano) on 24 original research papers using Mouse Mammary Tumor Virus-Like Virus and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal of papers, consisting of multiple choice, Likert-scale, multi-select, and free-text question types (77 question items across 24 papers). Agreement between (1) experts, and (2) experts and each LLM, was determined per question instance using novel scoring metrics. LLMs were assessed by comparing inter-expert and expert-LLM agreement score distributions to determine whether LLMs behaved as additional experts by either increasing or maintaining inter-expert agreement. Free-text responses were further evaluated qualitatively.
Across all question types, LLM responses aligned closely with expert assessments, with two models (GPT-5, GPT-5 Nano) achieving score distributions statistically indistinguishable from those of experts. Gemini models behaved similarly for most tasks but were significantly more lenient in applying microbial oncogenesis criteria, often over-attributing criteria fulfillment. Hallucinations were rare, although more frequent in smaller models (Gemini 2.5 Flash, GPT-5 Nano). Methodological appraisal and identification of contradictions within full-text papers were the most persistent areas of LLM vulnerability, however, the error rate could not be directly compared with experts.
Two LLMs (GPT-5, GPT-5 Nano) were indistinguishable from domain experts on structured domain research paper evaluation tasks. This evidence supports use of LLMs for automated systematic evidence synthesis. However, methodological appraisal tasks and contradiction identification in full-text papers remain weaknesses requiring further investigation, strengthening, and possibly multi-model strategies.
Kaela Kokkas, Hairong Wang, Richard Klein et al.· Frontiers in Cellular and In...· 0 citations
Oncology notes contain the richest clinical detail, yet they remain largely inaccessible at scale because extracting structured phenotypes requires either substantial language model infrastructure, curated training data, or cloud computing under regulatory constraints. We developed OncoRAG, combining ontology enrichment, knowledge graph construction, graph-diffusion reranking, and structured prompting with a locally deployed 14B-parameter language model without model weight fine-tuning. Applied to three cohorts—triple-negative breast cancer (TNBC; 104 patients, 42 features; primary development), recurrent high-grade glioma (RiCi; 191 patients, 19 features; cross-lingual and cross-disease evaluation with cohort-specific configuration), and MIMIC-IV (100 patients, 10 features; limited external evaluation on overlapping features)—OncoRAG achieved F1 scores of 0.80, 0.79, and 0.84, improving over direct large language model (LLM) prompting and naive retrieval-augmented generation (RAG) baselines by 0.19–0.22 and 0.17–0.19 F1, and outperforming direct prompting with a 5× larger 70B model by 0.09–0.10 F1. In an exploratory survival analysis (12 events), both feature sets showed close point estimates of the C-index (0.77 vs 0.76), but equivalence cannot be statistically confirmed given the limited event count. OncoRAG enables accurate clinical phenotyping from multilingual oncology notes using a locally deployable mid-size model, without model weight fine-tuning or external data sharing.
P. Salome, Maximilian Knoll, David Walz et al.· npj Digital Medicine· 0 citations
Summary Diagnosing rare diseases remains a major challenge due to limited clinical knowledge and the frequent absence of diagnostic criteria. We present a digital framework that leverages large language models and biomedical text embeddings to bridge this gap. By mapping Human Phenotype Ontology terms to a shared vector space with millions of PubMed abstracts and full-text articles, our method enables phenotype-driven semantic search and ranks literature relevant to patient symptoms, even without explicit disease mentions. Validated on OMIM-derived benchmarks and applied to RASopathies, including NF1, Noonan, and Costello syndromes, our approach retrieved expected findings, supporting differential diagnosis and research. The framework is implemented in an open-source Python package, py-semtools, and it can be integrated into clinical decision support systems or adapted to other ontologies and corpora. This work demonstrates how AI-driven informatics can enhance rare disease diagnosis and exemplifies the role of digital tools in transforming precision medicine and healthcare delivery.
Jesús Pérez-García, Federico García-Criado, F. Pazos et al.· iScience· 0 citations