Skip to content
Open access

Interpretable Small-Sample Deep Learning with Approximate Attribution Enhancement for Atomically Precise Gold Nanoclusters Synthesis

Aug 2026 · Molecules · Vol 31, pp. 2808 · 0 citations · 57 references
Medicine

TL;DR

The results position AAE as a local regularisation strategy for data-limited synthesis modelling rather than as a source of new chemical information.

Abstract

Small experimental datasets make synthesis-condition modelling particularly sensitive to overfitting and data leakage. Here, a graph convolutional neural network (GCNN) was coupled with approximate attribution enhancement (AAE) and evaluated on 54 gold-nanocluster synthesis records using a strict fold-local pipeline. All preprocessing, supervised representation learning, augmentation, model selection, and calibration were confined to each training fold, and metrics were calculated only on untouched original records. The strongest no-augmentation model achieved a Matthews correlation coefficient (MCC) of 0.729. With 20,000 fold-local AAE samples, Logistic Regression reached an accuracy of 0.864, an F1 score of 0.862, and an MCC of 0.758. The corresponding MCC values decreased to 0.579, 0.526, and 0.652 under leave-one-ligand-out, leave-one-publication-out, and leave-region-out evaluation, respectively, indicating that extrapolation beyond represented chemistry remained more difficult than random-fold interpolation. Feature, label-weight, augmentation-baseline, covariance, calibration, and stability analyses consistently identified ligand aromaticity, temperature, and HAuCl4 concentration as the most influential factors. An interpretable analysis of a 34-record aqueous Au25 subset further produced condition rules and a pH–temperature map, while nine literature formulations provided an external audit within the represented domain. The results position AAE as a local regularisation strategy for data-limited synthesis modelling rather than as a source of new chemical information.

Read PDF

Similar papers

Aug 2026

MCCX: Combining Monte Carlo, Convolutional Neural Networks and Boosting in Predicting Nanoparticle Viability

Modeling and prediction relying on learning from small experimental data sets is challenging, especially when the data set is insufficient, noisy or involves highly complex relationships. Leveraging Neural Networks (NNs), we developed predictive models capable of addressing such small data sets. We introduce MCCXcom...

Qiankun Mo, Xing-Fei Wei, Ravithree D. Senanayake et al. · 0 citations
Preprint Sep 2026

Pretraining and Distillation Matter More Than Architecture Family for Label-Free Single-Cell Classification

Choosing a deep learning architecture for label-free single-cell classification remains an open question, with microscopy benchmarks reporting conflicting conclusions about CNNs versus transformers. We present a controlled benchmark on LIVECell phase-contrast microscopy data using source-image-disjoint train/validation...

Philip Graemer, G. Di Caprio · 0 citations
Open access 2026

Cross-Domain Evaluation of Modern Deep Learning Architectures for Microscopic Diatom Classification

Automatic classification of microscopic images is important for water-quality monitoring and ecosystem-health assessment, but existing deep-learning solutions are usually validated on single-source datasets and overlook the domain shift that arises across laboratories. This study evaluates how five representative deep-...

U. Yurtsever · 0 citations
#machine learning Preprint Aug 2026

ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity Prediction

ToxLens is introduced, a reproducible multi-task graph-learning framework for 11 toxicity endpoints spanning Ames mutagenicity, acute oral toxicity, hERG inhibition, and Tox21 nuclear-receptor and stress-response assays and reveals substantial endpoint-specific variation in set efficiency and discrimination and calibra...

Magnus H. Strømme, A. D. de Sá, David B. Ascher · 0 citations
Sep 2026

Optimization-driven feature selection for enhanced skin cancer detection using deep learning

Improved robustness and explainability in lesion classification is demonstrated, supporting the applicability of the proposed method to real clinical settings and future research focuses on multimodal fusion and lightweight deployment on edge devices.

Amruta Thorat, Chaya Jadhav · 0 citations
#protein folding Preprint Sep 2026

MIRAGE: Measuring Interpolation and Redundancy in Affinity GEneralization

MIRAGE (Measuring Interpolation and Redundancy in Affinity GEneralization), a plug-in benchmark treating historical public family support (through 2019) as an explicit variable, applying a family-support axis to affinity and pose prediction via matched strata, family-disjoint controls, ligand-only baselines, and tempor...

Mehdi Yazdani-Jahromi, Sanjay Padhi, Ivan Garibay · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.