Enhancing Cardiovascular Risk Stratification with Interpretable Ensemble Learning: A Comparative Analysis and Explainable AI-Driven Insights
DOI:
https://doi.org/10.21512/commit.v20i2.13355Keywords:
Machine Learning, Ensemble Learning, Cardiovascular Disease Prediction, AdaBoost, Bagging, Explainable AI, SHapley Additive exPlanations (SHAP)Abstract
Cardiovascular disease (CVD) is a leading cause of global mortality, demanding accurate and interpretable risk stratification models for early intervention. However, the tradeoff between model performance and clinical interpretability remains a gap. Therefore, the research evaluates lightweight parametric ensemble models against complex black-box alternatives. A comparative analysis of Logistic Regression (LR) and Gaussian Naive Bayes (GNB) is conducted using a publicly available dataset of 1,000 patient records, enhanced with Bagging and AdaBoost ensemble methods. The data undergo standardized preprocessing to mitigate bias, and a fivefold stratified cross-validation protocol ensures model generalizability. The Bagging LR ensemble achieves the best predictive performance, with a mean accuracy of 0.966 (95% Confidence Interval (CI): [0.958, 0.974]) and a Receiver Operating Characteristic-Area Under the Curve (ROC-AUC) of 0.993 (95% CI: [0.990, 0.996]), highly competitive with state-of-the-art baselines including LightGBM and CatBoost. These results show that computationally efficient and interpretable ensembles provide a viable alternative to more complex models, particularly in data-constrained clinical settings. A comprehensive Explainable Artificial Intelligence (XAI) analysis using SHapley Additive exPlanations (SHAP) identifies clinically congruent drivers of prediction, the slope of the peak exercise ST segment, ST depression, and chest pain type. The findings highlight the potential of interpretable ensemble learning to develop dependable, deployable clinical decision support tools for integration into hospital or e-health systems for early CVD risk stratification. Further validation on larger, more diverse datasets is recommended to confirm broader clinical applicability.
References
[1] T. Adam et al., “The state of cardiac rehabilitation in Saudi Arabia: Barriers, facilitators, and policy implications,” Cureus, vol. 15, no. 11, pp. 1–11, 2023.
[2] K. E. Setiawan, A. Kurniawan, A. Chowanda, and D. Suhartono, “Clustering models for hospitals in Jakarta using Fuzzy C-Means and K-Means,” Procedia Computer Science, vol. 216, pp. 356–363, 2023.
[3] P. Wicaksono, P. Samuel, I. N. Alam, and S. M. Isa, “Dealing with imbalanced sleep apnea data using DCGAN,” Traitement du Signal, vol. 39, no. 5, pp. 1527–1536, 2022.
[4] Y. Q. Cai et al., “Pitfalls in developing machine learning models for predicting cardiovascular diseases: Challenge and solutions,” Journal of Medical Internet Research, vol. 26, pp. 1–21, 2024.
[5] S. Ouf and A. I. B. ElSeddawy, “A proposed paradigm for intelligent heart disease prediction system using data mining techniques,” Journal of Southwest Jiaotong University, vol. 56, no. 4, pp. 220–240, 2021.
[6] S. P. Praveen et al., “Enhanced feature selection and ensemble learning for cardiovascular disease prediction: Hybrid GOL2-2 T and adaptive boosted decision fusion with babysitting refinement,” Frontiers in Medicine, vol. 11, pp. 1–18, 2024.
[7] D. Asif, M. Bibi, M. S. Arif, and A. Mukheimer, “Enhancing heart disease prediction through ensemble learning techniques with hyperparameter optimization,” Algorithms, vol. 16, no. 6, pp. 1– 17, 2023.
[8] S. J. Lee et al., “Deep learning improves prediction of cardiovascular disease-related mortality and admission in patients with hypertension: Analysis of the Korean National Health Information Database,” Journal of Clinical Medicine, vol. 11, no. 22, pp. 1–12, 2022.
[9] A. Ogunpola, F. Saeed, S. Basurra, A. M. Albarrak, and S. N. Qasem, “Machine learning-based predictive models for detection of cardiovascular diseases,” Diagnostics, vol. 14, no. 2, pp. 1–19, 2024.
[10] M. M. Yaqoob, M. Nazir, M. A. Khan, S. Qureshi, and A. Al-Rasheed, “Hybrid classifier-based federated learning in health service providers for cardiovascular disease prediction,” Applied Sciences, vol. 13, no. 3, pp. 1–17, 2023.
[11] X. Gao, S. Alam, P. Shi, F. Dexter, and N. Kong, “Interpretable machine learning models for hospital readmission prediction: A two-step extracted regression tree approach,” BMC medical informatics and decision making, vol. 23, pp. 1–11, 2023.
[12] D. Y. Omkari and S. B. Shinde, “Opportunities and challenges of machine learning and deep learning techniques in cardiovascular disease prediction: A systematic review,” Journal of Biological Systems, vol. 31, no. 02, pp. 309–344, 2023.
[13] M. A. Naser, A. A. Majeed, M. Alsabah, T. R. Al-Shaikhli, and K. M. Kaky, “A review of machine learning’s role in cardiovascular disease prediction: recent advances and future challenges,” Algorithms, vol. 17, no. 2, pp. 1–33, 2024.
[14] P. Aryawibowo, A. F. Hidayanto, Y. M. Toemali, K. E. Setiawan, and A. A. S. Gunawan, “Intelligent monitoring and diagnosing capability in healthcare: Systematic literature review,” in 2023 International Conference on Information Management and Technology (ICIMTech). Malang, Indonesia: IEEE, Aug. 24–25, 2023, pp. 627–632.
[15] R. Katarya and S. K. Meena, “Machine learning techniques for heart disease prediction: A comparative study and analysis,” Health and Technology, vol. 11, no. 1, pp. 87–97, 2021.
[16] A. Eleyan and E. Alboghbaish, “Electrocardiogram signals classification using deep-learningbased incorporated convolutional neural network and long short-term memory framework,” Computers, vol. 13, no. 2, pp. 1–12, 2024.
[17] S. Xian et al., “Transformer patient embedding using electronic health records enables patient stratification and progression analysis,” npj Digital Medicine, vol. 8, no. 1, pp. 1–16, 2025.
[18] J. Dumlao, “Cardiovascular disease dataset,” 2023. [Online]. Available: https://www.kaggle.com/datasets/jocelyndumlao/cardiovascular-disease-dataset
[19] G. Ngo, R. Beard, and R. Chandra, “Evolutionary bagging for ensemble learning,” Neurocomputing, vol. 510, pp. 1–14, 2022.
[20] H. Jain, A. Khunteta, and S. Srivastava, “Churn prediction in telecommunication using logistic regression and logit boost,” Procedia Computer Science, vol. 167, pp. 101–112, 2020.
[21] V. Vangara, S. P. Vangara, and K. Thirupathur, “Opinion mining classification using naive bayes algorithm,” International Journal of Innovative Technology and Exploring Engineering (IJITEE), vol. 9, no. 5, pp. 495–498, 2020.
[22] S. Gonz´alez, S. Garc´ıa, J. Del Ser, L. Rokach, and F. Herrera, “A practical tutorial on bagging and boosting based ensembles for machine learning: Algorithms, software tools, performance study, practical perspectives and opportunities,” Information Fusion, vol. 64, pp. 205–237, 2020.
[23] P. P. Pande, “Special issue on benchmarking machine learning systems and applications,” IEEE Design & Test, vol. 39, no. 3, 2022.
[24] S. A. Hicks and Others, “On evaluation metrics for medical applications of artificial intelligence,” Scientific Reports, vol. 12, pp. 1–9, 2022.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Panji Arisaputra, Pandu Wicaksono, Karli Eka Setiawan, Anindhita Dewabharata

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
a. Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License - Share Alike that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
b. Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
c. Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
Â
USER RIGHTS
All articles published Open Access will be immediately and permanently free for everyone to read and download. We are continuously working with our author communities to select the best choice of license options, currently being defined for this journal as follows: Creative Commons Attribution-Share Alike (CC BY-SA)

















