Skip to content
Open access

TriFusion-PCGNet: A Hybrid Multi-Branch Deep Learning Network for Energy-Aware Normal/Abnormal Heart Sound Classification

Aug 2026 · Life · Vol 16, pp. 1417 · 0 citations · 36 references

TL;DR

The substantially lower subject-disjoint performance of TriFusion-PCGNet indicates that the record-level results should not be interpreted as evidence of patient-independent generalization, and further validation on independent subject-annotated datasets is required.

Abstract

Background/Objectives: Cardiovascular diseases remain a leading cause of death worldwide, and early diagnosis is essential. Phonocardiogram (PCG) signals provide a non-invasive and inexpensive means of cardiac assessment. However, accurate interpretation of PCG signals requires clinical expertise. In this paper, we propose TriFusion-PCGNet, a hybrid multi-branch deep learning framework for automated heart sound classification. Methods: The proposed method was evaluated on the PhysioNet/Computing in Cardiology Challenge 2016 dataset. Three complementary PCG representations—the raw signal, band-pass filtered signal, and Hilbert envelope signal—were processed through dedicated feature extraction branches using residual convolutional networks, dilated residual convolutions, and a BiGRU with a temporal attention mechanism. The extracted features were fused and classified using fully connected layers. Performance was evaluated using the original train–validation split and stratified 10-fold cross-validation. Results: Under a strictly separated record-level 10-fold cross-validation protocol, TriFusion-PCGNet achieved 90.87% mean accuracy and 95.34% AUC, with outer test folds isolated from training and model selection. Subject-level analysis showed substantially lower performance, emphasizing the importance of subject-disjoint evaluation for patient-level generalization. Conclusions: TriFusion-PCGNet integrates complementary PCG representations for automated normal/abnormal heart sound classification. However, the substantially lower subject-disjoint performance indicates that the record-level results should not be interpreted as evidence of patient-independent generalization, and further validation on independent subject-annotated datasets is required.

Read PDF

Similar papers

Open access Aug 2026

From convolutions to transformers: a comparative analysis of deep learning models for heart sound classification

Cardiovascular diseases remain one of the primary causes of death worldwide, thus generating an ever-continuing market demand for authentic and non-invasive diagnostic tools. The current research focuses on the automatic classification of heart sound signal recordings (phonocardiograms, PCGs) with deep learning models to facilitate detection of cardiac abnormalities. Three deep learning approaches were compared in a common experimental framework: a baseline Convolutional Neural Network (CNN), an Attention Augmented CNN (Attention-CNN), and Audio Spectrogram Transformer (AST). Experiments were performed on the PhysioNet/CinC Challenge 2016 data set using Mel-spectrogram representations. The performance of all the architectures was evaluated using a 5-fold cross-validation protocol with Area Under the Receiver Operating Characteristic Curve (AUC), Accuracy, and F1 metrics. Training used the Adam optimizer with binary cross-entropy loss, and record-level predictions were obtained by aggregating segment-level probabilities for each recording. Attention-CNN got the best results with an AUC of 0.9711 and an F1-score of 0.7679, beating the baseline CNN (AUC = 0.9636) and AST (AUC = 0.9243). The use of channel attention resulted in better feature discrimination and fewer false positive predictions. While AST demonstrated the ability to capture global contextual features through self-attention, its performance was constrained by computational resources and limited training period. The findings suggest that Attention-CNN is a preferable balanced relationship between performance and computational requirements for PCG-based heart sound classification. However, the study has limitations in terms of data size and availability of computational resources, especially for transformer models. Future work will continue with larger scale training and validation to further test model generalizability in clinical and telemedicine practice.

Bollapalli Althaph, Devyani Sah, Lasya Reddy et al. · 0 citations
Open access Aug 2026

STFT-Based Multiclass Heart Sound Classification Using BiLSTM and CNN-BiLSTM Models

Cardiovascular diseases require early and reliable screening because manual auscultation may be affected by noise, subjective interpretation, and inter-observer variability. This study aimed to develop an STFT-based deep learning framework for multiclass phonocardiogram (PCG) classification. The proposed framework was designed to provide a reproducible evaluation procedure by combining standardized preprocessing, time–frequency feature extraction, and deep learning-based classification under the same experimental conditions. Unlike approaches that may evaluate segmented signals without clearly preserving recording-level separation, this study emphasizes a leakage-free splitting strategy to reduce the risk of overestimated performance and to provide a more reliable assessment of model generalization. The Yaseen PCG dataset, consisting of 1000 recordings from five classes (AS, MR, MS, MVP, and Normal), was divided using a leakage-free recording-level split before segmentation and spectrogram generation. After preprocessing, 2-second PCG segments with 50% overlap were converted into 128 × 128 STFT spectrograms and classified using BiLSTM and CNN-BiLSTM models. Both models were trained and tested using the same dataset split, preprocessing pipeline, and evaluation metrics, including accuracy, precision, recall, F1-score, specificity, and confusion matrices. The BiLSTM model achieved 92.36% accuracy in the final independent test run, while the CNN-BiLSTM model achieved 95.83%. Across three repeated runs, BiLSTM achieved 93.85% ± 1.47%, whereas CNN-BiLSTM achieved 95.94% ± 0.71%. These results show that CNN-BiLSTM provides higher and more stable classification performance for five-class PCG classification, while BiLSTM remains a simpler alternative for lightweight implementation. Overall, the proposed STFT-based framework provides a reliable approach for automated heart sound classification and may support future computer-aided cardiac screening applications.

N. S., E. A. Hussein, L. Abdul-Rahaim · 0 citations
Open access Aug 2026

HSD-Net: a dual-branch CNN-BiLSTM network with hybrid cepstral fusion for heart sound classification.

Introduction Cardiovascular diseases (CVDs) represent a major global health threat, making early detection crucial. While cardiac auscultation is cost-effective, its reliance on clinical expertise leads to significant variability in diagnostic accuracy. Existing automated heart sound classification methods suffer from inadequate feature representation and limited model performance due to constraints imposed by the randomness, nonstationarity, and complex acoustic environment of phonocardiogram (PCG) signals. Methods In this paper, an improved dual-branch CNN-BiLSTM network with hybrid cepstral fusion, named HSD-Net, is proposed for automated heart sound classification. To overcome feature representation limitations, we introduce a hybrid cepstral fusion feature set that combines Mel-Frequency Cepstral Coefficients (MFCC) and Inverse Mel-Frequency Cepstral Coefficients (IMFCC) with their first-order delta coefficients, capturing complementary low-frequency and high-frequency acoustic information. This feature set is processed by a dedicated dual-branch architecture: a convolutional neural network (CNN) branch extracts localized spectrotemporal patterns, while a bidirectional Long Short-Term Memory (BiLSTM) branch models long-range temporal dependencies. A Squeeze-and-Excitation (SE) attention mechanism is integrated to adaptively recalibrate feature importance. Results Extensive evaluations on two public datasets demonstrate that the proposed HSD-Net achieves superior performance compared with existing methods. On the PhysioNet/Computing in Cardiology Challenge 2016 dataset, the model achieved an accuracy of 96.25%, a sensitivity of 95.85%, and a specificity of 96.50%. On the additional five-class Yaseen dataset, the model achieved an average accuracy of 99.32% in valvular heart disease classification. Ablation studies quantitatively confirmed that both the hybrid features and the dual-branch architecture contribute significantly to the overall performance. Discussion The proposed HSD-Net provides an effective framework for automated heart sound analysis. This work offers a reliable technical foundation for developing clinical decision-support systems, with significant potential for application in early screening and telemedicine.

Caijian Hua, Ye Tian, Liuying Li et al. · 0 citations
Conference Jul 2026

An Echocariogram based Heart Disease using Deep Learning Model

Heart disease is significant health burden and a major cause of death worldwide, hence the need to develop automated diagnostic systems with correctness and computational efficiency in order to rescue lives through early clinical intervention. This paper introduces a fine-grained deep transfer learning architecture to classify cardiac disease images into multiple classes using two large convolutional backbones ResNet50 and DenseNet121, which are trained systematically with AdamW, stochastic gradient descent (SGD), and Lion under a single experimental environment. The suggested pipeline combines standardized image preprocessing, stratified division of data, transfer learning hierarchical feature extraction based on the feature, hyper parameter convergence analysis, and 50-epoch supervised fine-tuning, and then thorough performance and efficiency assessment. The experimental findings indicate the stable optimization dynamic in all six backbone optimizer setups, with the most consistent validation convergence in ResNet50 with Lion optimizer. The most successful configuration obtained nearly 93.9% accuracy of validation, precision, recall, and F1-score and a robust ROC-AUC of 0.995, which verified a high inter-class separability and stable threshold-independent result. The loss and validation accuracy curves also show that the convergence is quick and the ability to generalize is high. In terms of deployment, DenseNet121 was found to have much lower architectural complexity (8.0M parameters) and reduced inference latency (8.4 ms/image) than ResNet50 (25.6M parameters, 12.8 ms/image), and yet have a competitive level of classification. The graph driven analysis also shows that optimizer selection mostly influences the smoothness of convergence, predictive calibration, but backbone architecture influences a tradeoff between representational richness and computational efficiency. On the whole, the suggested framework is a clinically applicable and deployment-focused solution to smart computer aided diagnosis of cardiac diseases.

Sarangam Kodati, Nadimpally Nutesh Goud · 0 citations
Open access Aug 2026

Explainable hybrid deep learning framework with Grad-CAM for heartbeat-level arrhythmia classification

Cardiac arrhythmia, a common sign of cardiovascular disease, is a leading cause of global mortality and morbidity. Timely and accurate detection of electrocardiogram (ECG) cardiac arrhythmia is essential for effective clinical intervention. Here, we present an explainable hybrid deep learning framework for automated ECG cardiac arrhythmia classification. The proposed model integrates one-dimensional convolutional neural network (1D-CNN) for local morphological features, a gated recurrent unit (GRU) to extract temporal dependencies in sequential ECG signals, and the channel attention technique to emphasize clinically important patterns. In addition, a gradient-weighted class activation mapping (Grad-CAM) module is integrated to enhance interpretability by highlighting complex portions of the ECG signals that influence model decision making. The proposed framework is evaluated on three benchmark datasets: Massachusetts Institute of Technology–Beth Israel Hospital (MIT-BIH) Arrhythmia, St. Petersburg INCART, and the MIT-BIH Supraventricular Arrhythmia Database (SVDB). The model achieves classification accuracy values of 99.69%, 99.73%, and 98.77% along with macro-F1 scores of 95.75%, 97.21%, and 94.58% and specificity values of 99.53%, 99.54%, and 98.63%, respectively. The experimental results demonstrated performance improvement of approximately 2%–5% over recent state-of-the-art techniques in terms of key performance metrics. These findings indicate that the proposed framework effectively integrates local morphological feature extraction with temporal modeling, providing a robust and scalable solution for ECG arrhythmia classification. Moreover, its computational efficiency and its use of single-lead ECG signals make it suitable for real-time deployment in wearable devices, remote monitoring systems, and resource-limited clinical settings.

S. Sundaramoorthy, Govardhan Karunanidhi · 0 citations