Artificial intelligence in cellular senescence research: a systematic review assessing methodological quality and reporting standards using PROBAST + AI and TRIPOD + AI
Cellular senescence is a fundamental mechanism of biological ageing that has emerged as a critical target for therapeutic intervention in age related diseases. The coalesce of artificial intelligence and senescence research provides unprecedented opportunities in advancing our knowledge and treatment approaches. This systematic review study addresses the gap across diverse AI model and the heterogeneity in senescence, by conducting the extensive evaluation of the performance outcomes and methodological rigor of AI models using Prediction model Risk of Bias Assessment Tool + Artificial Intelligence (PROBAST + AI) and reporting completeness using Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis + Artificial Intelligence extension (TRIPOD + AI), providing the insights into AI models robustness and generalizability. For the articles, released between 2019 to 2025, across major databases 18 eligible studies was used for review, following PRISMA guideline. Quality and applicability were assessed by PROBAST + AI (4 domain) and reporting via TRIPOD + AI (27 items). The quantitative synthesis indicates that deep learning architectures, especially Convolutional Neural Networks (CNNs), are dominant which appeared in about 50% of the studies. These CNNs consistently outperform traditional machine learning methods in the analysis of morphological heterogeneity. While reported performance metrics were high, with accuracy ranging from 83.55% to 99.79%, the PROBAST + AI assessment indicates a high risk of bias in 83.33% (15/18) of studies, primarily driven by Analysis domain due to improper data splitting (data leakage) and lack of external validation. As well as adherence to TRIPOD + AI reporting standards was suboptimal with average of 62% ‘YES’; notably, with the major gap in 0% of studies pre-registered a protocol and only 44.4% made analytical code publicly available, severely limiting reproducibility. Evidently AI demonstrates immense potential to accelerate biomarker discovery and senolytic drug screening, particularly through label-free morphological analysis by DL models, despite of high quality concern and poor reproducibility limit reliability; also standardization, shared benchmarks, multi-omics integration, and explainable AI are essential concerns for clinical translation in aging research.
GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations of such models.
Xiaotian Zhang, Chun-yan Li, Yi Zong et al.· arXiv.org· 216 citations· ⚡17
Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.
Jinhe Bi, Yifan Wang, Danqi Yan et al.· arXiv.org· 73 citations· ⚡4
The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.
Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson· EUROMICRO Conference on Soft...· 64 citations· ⚡6
In this paper, we present a novel approach to improving software quality and efficiency through a Large Language Model (LLM)-based model designed to review code and identify potential issues. Our proposed LLM-based AI agent model is trained on large code repositories. This training includes code reviews, bug reports, and documentation of best practices. It aims to detect code smells, identify potential bugs, provide suggestions for improvement, and optimize the code. Unlike traditional static code analysis tools, our LLM-based AI agent has the ability to predict future potential risks in the code. This supports a dual goal of improving code quality and enhancing developer education by encouraging a deeper understanding of best practices and efficient coding techniques. Furthermore, we explore the model's effectiveness in suggesting improvements that significantly reduce post-release bugs and enhance code review processes, as evidenced by an analysis of developer sentiment toward LLM feedback. For future work, we aim to assess the accuracy and efficiency of LLM-generated documentation updates in comparison to manual methods. This will involve an empirical study focusing on manually conducted code reviews to identify code smells and bugs, alongside an evaluation of best practice documentation, augmented by insights from developer discussions and code reviews. Our goal is to not only refine the accuracy of our LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.
Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al.· arXiv.org· 62 citations· ⚡3
This paper designs Markov decision processes (MDPs) for different combinatorial problems and proposes to train conditional GFlowNets to sample from the solution space and demonstrates that GFlowNet policies can efficiently find high-quality solutions.
Dinghuai Zhang, H. Dai, Esmeralda S. Whitammer et al.· Advances in Neural Informati...· 59 citations· ⚡8
An empirical study on the current state of practice in artificial intelligence ethics is conducted by means of a multiple case study of five case companies, which indicates a gap between research and practice in the area.
Ville Vakkuri, Kai-Kristian Kemell, Joni Kultanen et al.· arXiv.org· 56 citations· ⚡6
The professor of physics and inaugural director of the NSF AI Institute for Artificial Intelligence and Fundamental Interactions will lead LNS and continue his research in particle physics.