Aug 2026· Molecules· Vol 31, pp. 2808· 0 citations· 57 references
Medicine
TL;DR
The results position AAE as a local regularisation strategy for data-limited synthesis modelling rather than as a source of new chemical information.
Abstract
Small experimental datasets make synthesis-condition modelling particularly sensitive to overfitting and data leakage. Here, a graph convolutional neural network (GCNN) was coupled with approximate attribution enhancement (AAE) and evaluated on 54 gold-nanocluster synthesis records using a strict fold-local pipeline. All preprocessing, supervised representation learning, augmentation, model selection, and calibration were confined to each training fold, and metrics were calculated only on untouched original records. The strongest no-augmentation model achieved a Matthews correlation coefficient (MCC) of 0.729. With 20,000 fold-local AAE samples, Logistic Regression reached an accuracy of 0.864, an F1 score of 0.862, and an MCC of 0.758. The corresponding MCC values decreased to 0.579, 0.526, and 0.652 under leave-one-ligand-out, leave-one-publication-out, and leave-region-out evaluation, respectively, indicating that extrapolation beyond represented chemistry remained more difficult than random-fold interpolation. Feature, label-weight, augmentation-baseline, covariance, calibration, and stability analyses consistently identified ligand aromaticity, temperature, and HAuCl4 concentration as the most influential factors. An interpretable analysis of a 34-record aqueous Au25 subset further produced condition rules and a pH–temperature map, while nine literature formulations provided an external audit within the represented domain. The results position AAE as a local regularisation strategy for data-limited synthesis modelling rather than as a source of new chemical information.
Modeling and prediction relying on learning from small experimental data sets is challenging, especially when the data set is insufficient, noisy or involves highly complex relationships. Leveraging Neural Networks (NNs), we developed predictive models capable of addressing such small data sets. We introduce MCCXcom...
Qiankun Mo, Xing-Fei Wei, Ravithree D. Senanayake et al.· ACS Applied Engineering Mate...· 0 citations
Choosing a deep learning architecture for label-free single-cell classification remains an open question, with microscopy benchmarks reporting conflicting conclusions about CNNs versus transformers. We present a controlled benchmark on LIVECell phase-contrast microscopy data using source-image-disjoint train/validation...
Automatic classification of microscopic images is important for water-quality monitoring and ecosystem-health assessment, but existing deep-learning solutions are usually validated on single-source datasets and overlook the domain shift that arises across laboratories. This study evaluates how five representative deep-...
ToxLens is introduced, a reproducible multi-task graph-learning framework for 11 toxicity endpoints spanning Ames mutagenicity, acute oral toxicity, hERG inhibition, and Tox21 nuclear-receptor and stress-response assays and reveals substantial endpoint-specific variation in set efficiency and discrimination and calibra...
Magnus H. Strømme, A. D. de Sá, David B. Ascher· 0 citations
Improved robustness and explainability in lesion classification is demonstrated, supporting the applicability of the proposed method to real clinical settings and future research focuses on multimodal fusion and lightweight deployment on edge devices.
MIRAGE (Measuring Interpolation and Redundancy in Affinity GEneralization), a plug-in benchmark treating historical public family support (through 2019) as an explicit variable, applying a family-support axis to affinity and pose prediction via matched strata, family-disjoint controls, ligand-only baselines, and tempor...
Mehdi Yazdani-Jahromi, Sanjay Padhi, Ivan Garibay· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.