Experiments show that PWAV generally improves over classical fingerprint descriptors within learned models and achieves competitive performance relative to established external baselines on several endpoints, positioning PWAV as a competitive and chemically transparent component for hybrid molecular property prediction, rather than as a replacement for domain-specific benchmark systems.
Abstract
Molecular property prediction is central to cheminformatics and environmental chemistry, where accurate modeling of physicochemical properties supports risk assessment and molecular design. Classical descriptors and recent advances such as ChemBERTa have enabled learning chemically contextual representations directly from SMILES, while the integration of structured descriptors with transformer-based embeddings offers a promising pathway toward accurate and interpretable prediction. In this study, we introduce Path-Weighted Atom Vectors (PWAVs), a descriptor family that captures atom-level, environment-aware structural information. We evaluate PWAV both as a standalone representation and in combination with ChemBERTa embeddings through a gated fusion architecture incorporating modality dropout, FiLM conditioning, and auxiliary supervision. Experiments on six physicochemical property datasets ( log P, log S, log BCF, boiling point, melting point, and vapor pressure) show that PWAV generally improves over classical fingerprint descriptors within learned models and achieves competitive performance relative to established external baselines on several endpoints. The strongest gains are observed for boiling point, aqueous solubility, and partition coefficient prediction, where descriptor-embedding fusion yields the best results among the learned models considered. Ablation analyses demonstrate that PWAV contributes complementary structural information beyond SMILES-only ChemBERTa representations, while SHapley Additive exPlanations-based interpretability shows that predictive signal is concentrated within a compact subset of features, enabling an efficient reduced representation (PWAV-64). Nested cross-validation further confirms the robustness of PWAV within the XGBoost framework. Overall, PWAV provides a compact, interpretable, and extensible descriptor framework that integrates effectively with modern representation-learning approaches. These results position PWAV as a competitive and chemically transparent component for hybrid molecular property prediction, rather than as a replacement for domain-specific benchmark systems.
This review provides a systematic overview of recent advances in SSL-based molecular property prediction and analyzes how multimodal molecular representation learning by integrating sequence, graph, three-dimensional structure, and textual information can improve the quality and expressiveness of molecular representations.
Shuning Yang, Lei Deng· Journal of Chemical Informat...· 0 citations
FragBERTa is introduced, a molecular fragment-aware transformer-based representation language model pretrained using masked language modeling on Sequential Attachment-based Fragment Embedding (SAFE) representations, suggesting that fragment-based string representations offer advantages over atom-level representations for scaffold-sensitive and interaction-driven tasks.
Trimole-Hybrid is presented, a task-wise multimodal framework that addresses ADMET heterogeneity by selecting or combining predictors built from complementary molecular representations, and shows sensitivity to changes in essential functional motifs, suggesting its ability to capture ADMET-relevant molecular substructures.
A novel Dual-Attention Multimodal framework for Graphs and Sequence-based representations, so-called DAM-GS, which provides a promising solution for molecular property prediction with broad applications in drug discovery and computational molecular science.
Bay Van Nguyen, Vinh Truong, Ha Duong Thi Hong et al.· Journal of Chemical Informat...· 0 citations
The SBMR-CNN model demonstrates highly competitive accuracy, outperforming the CM, Uni-Mol+, and MPNN-2D benchmarks, while closely approaching the performance of the more computationally intensive MPNN-3D and SOAP descriptors, as well as the RF-MF model.
Abdulaziz W. Alherz, C. Tezak, Mohammed S. Alhajeri· Industrial & Engineering...· 0 citations
Accurate prediction of physicochemical properties is increasingly limited by an information ceiling of structure-only molecular descriptors. Here, predicted 1H|13C NMR chemical shifts are transformed into fixed-length NMR vectors and concatenated with ECFP4 to form the hybrid spectral–structural representation SpectraPRINTS, enabling direct evaluation of representational complementarity across logP, logS, and logD (pH 2.6, 7.4, and 10.5), as well as the most acidic and most basic pKas. With a fixed learning protocol, SpectraPRINT reduces error for lipophilicity- and solubility-related end points (up to 39% lower RMSE vs ECFP4), while no systematic gain is observed for the most acidic and most basic macroscopic pK a end points. The workflow is released as NMR-AI, a freely accessible web platform integrating NMR spectra prediction, descriptor construction, and property prediction, enabling interactive use and independent validation. The NMR-AI platform is accessible at https://cheminformaticsportal.if-pan.krakow.pl/.
Wojciech Pietruś, Arkadiusz Leniak, R. Kurczab· Journal of Chemical Informat...· 0 citations