Integrating Molecular Design and Machine Learning for Predictive Modeling of Novel DPP-IV Inhibitors in Diabetes Mellitus
Abstract
The current investigation focuses on the creation of a supervised machine learning model targeting dipeptidyl peptidase-IV(DPP-IV) associated with diabetes mellitus (DM) for expedited drug discovery. A Random Forest regression model was built utilizing molecular fingerprints to anticipate biological activity (pIC₅₀) using a data set comprising of 4267 bioactive molecules. Statistical parameters were used to evaluate model performance using coefficient of determination (R²) and mean squared error (MSE), demonstrating strong predictive power and robustness. Additionally, the model was validated using visualization techniques such as true vs. predicted bioactivity plots and feature attribution plot. These recognize specific molecular features contributing to bioactivity. According to the ML model, 80 novel molecules were substantially designed. Digital Screening with a trained model suggested the path for13 promising candidates. 4 compounds demonstrated high predicted activity having higher than 8 pIC₅₀ values (8.43,8.42,8.08). It indicates robust potential as lead molecules. Machine learning integration with rational drug design signifies the categorization of potent drug molecules. The proposed findings contribute a strong foundation for the discovery of novel drug candidates against diabetes mellitus minimizing the expenditure and time needed for therapeutic discovery.