Artificial Intelligence and Multimodal Data Approaches for Diagnosis, Prognosis, and Treatment Decision Support in Diffuse Large B-Cell Lymphoma and Acute Myeloid Leukaemia: A Systematic Review
Abstract
Aims: To synthesise and critically appraise evidence on artificial intelligence (AI) and multimodal data approaches for diagnosis, prognosis, and treatment decision support in diffuse large B-cell lymphoma (DLBCL) and acute myeloid leukaemia (AML), and to compare how disease biology shapes model design and clinical readiness. Study Design: Systematic review with narrative synthesis. Place and Duration of Study: PubMed/MEDLINE, Embase, Scopus, Web of Science Core Collection, IEEE Xplore, and the Cochrane Library were searched from database inception to 29 July 2026. Methodology: English-language original studies involving human participants, clinical images, human-derived samples, or patient-derived datasets were eligible. Two reviewers independently screened records, assessed full texts, and extracted study characteristics, data modalities, analytical methods, validation approaches, and performance measures. PROBAST+AI domains were used to guide appraisal of eligible patient-level prediction models; exploratory multi-omics, molecular-subtyping, biomarker-discovery, and segmentation studies underwent structured descriptive appraisal. TRIPOD+AI informed assessment of reporting completeness. Owing to substantial clinical and methodological heterogeneity, findings were synthesised narratively. Results: Twenty-eight studies were included: 22 on DLBCL and six on AML. DLBCL research was dominated by PET/CT radiomics, deep imaging, digital pathology, and imaging-clinical-molecular fusion. Internally evaluated diagnostic models reported area under the curve values as high as 0.999, whereas externally validated prognostic models achieved values of 0.66-0.71 and showed inconsistent incremental improvement over the International Prognostic Index. AML studies mainly integrated genomic, transcriptomic, epigenomic, single-cell, proteomic, metabolic, and functional drug-response data. Prognostic models reported concordance indices of 0.72-0.81 and time-dependent area under the curve values of 0.795-0.899. The most frequent methodological concerns were retrospective sampling, high-dimensional modelling in modest cohorts, incomplete calibration reporting, and limited independent validation. Treatment-response models in both diseases remained retrospective or exploratory, and none demonstrated improved outcomes through AI-guided treatment allocation. Conclusion: AI and multimodal data approaches show greatest maturity for prognostic stratification, but diagnostic replacement and treatment selection remain unproven. Prospective, multicentre clinical-impact validation is the principal requirement for translation into routine care.