Enhancing lung cancer screening accuracy through a dual-pipeline deep learning and machine learning framework for integrated risk assessment
Lung cancer is a leading cause of death worldwide, with delayed or misdiagnosis further hindering professionals' ability to lower lung cancer-related mortality rates. Misdiagnosis could look like false positive results, where benign lung diseases are mistaken for malignant lung carcinomas, or more life-threatening false negative results, where harmful eosinophilic pneumonia is misinterpreted as a benign lung hamartoma. The objective of this study consists of the following goals: (1) to develop CNN-based classification models for lung nodule analysis, (2) to compare the performance of multiple deep learning models under consistent controlled environments, (3) to evaluate and present the interpretability of the models through a clinically grounded probability equation, and (4) to propose an integrated risk-scoring framework that combines the imaging and clinical pipelines into a single interpretable measure of lung cancer risk. However this framework is only theoretical as imaging and clinical models were developed using separate datasets; future studies should note that a practical development will require validation on a unified multimodel patient cohort. In regard to the creation of the models, the deep learning convolutional neural network, which utilized a YOLOv8n backbone trained on the IQ-OTH/NCCD dataset of 649 annotated CT images, achieved a mean average precision (mAP50) of 0.687 when evaluated against expert-verified annotations. The machine learning branch utilizes a structured clinical dataset of 309 patient records with symptom- and risk-based features, including age, smoking status, chronic disease, fatigue, wheezing, and shortness of breath, all evaluated under standardized conditions. These machine learning models, showed similarly promising results, with the Logistic Regression (Log) achieving the highest mean accuracy of 0.929, while CatBoost (CB) demonstrated the strongest overall balance with a Mean F1 score of 0.953 at optimal hyperparameter settings of 1,000 iterations and a tree depth of 8. Interpretability is further enhanced through a clinically grounded cancer probability equation. This framework that showcases both deep and machine learning models will offer insight into the effectiveness of incorporating artificial intelligence in the screen accuracy of lung cancer classification.