Skip to content
Open access

MLHNet-Lung: an attention-guided multi-level CNN–transformer fusion framework with CBAM and GeM pooling for explainable multiclass lung CT image classification

Aug 2026 · Frontiers in Medicine · Vol 13 · 0 citations · 68 references
Medicine

TL;DR

The proposed MLHNet-Lung framework for image-level classification of adenocarcinoma, large cell carcinoma, normal, and squamous cell carcinoma CT slices demonstrates the value of jointly modeling complementary intermediate and high-level CT representations; however, patient-level multicenter validation and clinically verified labels remain necessary before diagnostic application.

Abstract

Accurate differentiation of lung cancer subtypes from computed tomography (CT) images remains challenging because malignant classes may exhibit overlapping radiographic characteristics, while conventional convolutional neural networks frequently rely only on final-layer semantic representations and may underuse informative intermediate features. This study proposes MLHNet-Lung, an attention-guided multi-level CNN–Transformer framework for image-level classification of adenocarcinoma, large cell carcinoma, normal, and squamous cell carcinoma CT slices. The study-specific novelty lies in the coordinated integration of intermediate block-4 and final EfficientNet-B0 feature maps, their independent refinement through Convolutional Block Attention Modules, projection into a shared 256-dimensional space, adaptive generalized mean pooling, compact three-token Transformer interaction, and three-source feature fusion. The experiments used 7,800 CT images, comprising 3,900 unique original images and 3,900 training-only augmented copies, divided into 5,200 training, 1,300 validation, and 1,300 independent test images. Validation and test sets contained only unique, non-augmented originals, and duplicate screening was applied to reduce image-level leakage. On the independent test set, MLHNet-Lung achieved an accuracy of 0.9415, a macro F1-score of 0.9414, and a macro ROC-AUC of 0.9821. Under the same experimental protocol, the proposed framework improved macro F1-score by 0.0104 over the strongest competing baseline and by 0.0175 over the standard EfficientNet-B0 classifier. The improvement over the strongest baseline was statistically significant according to McNemar’s test (p=0.028). Grad-CAM analysis further indicated that the model primarily emphasized pulmonary regions associated with its predictions. These findings demonstrate the value of jointly modeling complementary intermediate and high-level CT representations; however, patient-level multicenter validation and clinically verified labels remain necessary before diagnostic application.

Read PDF

Similar papers

Open access Aug 2026

MSTF-Net: A Multi-Scale Transformer-Guided Feature Fusion Framework for Accurate Lung Cancer Detection from CT Images

Lung cancer is one of the most common causes of cancer death worldwide, and early and accurate diagnosis is crucial to upsurge patient survival. Deep learning-based computer-aided diagnosis systems have demonstrated promising results, however they are difficult to recognise fine-grained local lesion traits and long-ran...

Pramod Kumar · 0 citations
Aug 2026

Lesion-Aware Multi-Scale CNN–Swin Transformer with Adaptive Attention Fusion for Explainable COVID-19 Detection from Chest CT Images

The rapid identification of COVID-19 from chest computed tomography (CT) images remains challenging due to the diverse appearance, distribution, and scale of pulmonary abnormalities. Conventional convolutional neural networks (CNNs) can capture detailed local patterns but may have limited ability to model long-range sp...

Vediya Raghuvanshi, P. Deore · 0 citations
Open access Sep 2026

Multi-view attention-based deep learning for benign–malignant classification of pulmonary nodules

Accurate classification of pulmonary nodules is essential for early lung cancer diagnosis, yet the limited spatial information provided by a single computed tomography slice may hinder the characterization of complex nodule morphology. Although multi-view strategies can capture complementary anatomical information, con...

Li-Jun Zhang, De-Bing Zhuo, Xiao-Lu Wu et al. · 0 citations
Conference Aug 2026

An Explainable CBAM Enhanced DenseNet121 Framework for Multi-Class Lung Cancer Classification Using CT Scans

Due to its late identification and challenging diagnosis, lung cancer continues to be one of the top causes of death for cancer patients globally, positioning it as one of the most critical concerns. Timely identification of cancerous nodules is essential for enhancing the patient’s survival likelihood CT image analysi...

S. Jegadeesan, S. Matheswaran, R. Palanivelrajan · 0 citations
Open access Sep 2026

Dual-Scale Hybrid Concept Bottleneck Network for Explainable 3D Lung Nodule Malignancy Classification in CT Imaging

Accurate differentiation of benign and malignant lung nodules in computed tomography (CT) is important for early lung cancer diagnosis and reliable clinical decision-making. Many existing deep learning methods emphasize either nodule-centred morphology or broader anatomical context and provide limited insight into the...

Ahmad Ali, Muhammad Aksam Iftikhar, Ghulam Farooque et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.