Volume 2, Issue 1 - July 2026
The reliability of machine learning-based heart disease prediction depends not only on the choice of predictive models but also on the quality and preparation of the underlying clinical data. Missing values, inconsistent variable representations, and differences in feature scales can affect the reliability of predictive models and their ability to distinguish between patients with and without heart disease. This study investigated the impact of systematic data pre-processing and missing-value handling on machine learning-based heart disease prediction. The study employed a hybrid methodological framework combining the Cross-Industry Standard Process for Data Mining (CRISP-DM) and Object-Oriented Methodology (OOM). CRISP-DM guided the data-driven process from data understanding and preparation to modelling and evaluation, while OOM supported the modular design and implementation of the system components. The study utilized the Cleveland Heart Disease dataset obtained from the University of California Irvine (UCI) Machine Learning Repository. A structured pre-processing pipeline involving data quality assessment, missing-value treatment using K-Nearest Neighbours (KNN) imputation, categorical variable encoding, and feature normalization using Standard Scalar was applied to prepare the dataset for predictive modelling. Logistic Regression, Random Forest, and Boost models were evaluated using accuracy, precision, recall, F1-score, and ROC-AUC. The ROC-AUC results showed that Logistic Regression achieved the highest discriminative performance with an AUC of 0.591, followed by Boost with 0.566, while Random Forest recorded the lowest AUC of 0.514, indicating performance close to random classification. These findings demonstrate that classification accuracy alone may not adequately reflect a model's discriminative capability and highlight the importance of using multiple evaluation metrics in clinical prediction. The study provides a structured and reproducible pre-processing and modelling framework for preparing clinical heart disease data and supports more rigorous evaluation of machine learning approaches for heart disease prediction.
Heart disease prediction, machine learning, data pre-processing, missing-value handling, KNN imputation, CRISP-DM, Object-Oriented Methodology, ROC-AUC, clinical data
Rebecca Nneka Oguzor, Dr. Chinagolum Ituma, Favour Ngozi Okoro, "Impact of Data Preprocessing and Missing-Value Handling on Machine Learning-Based Heart Disease Prediction", Cosmo Research & Science International Journal, vol. Jul-25, no. 1, pp. 607-627, 2026.
Rebecca Nneka Oguzor, Dr. Chinagolum Ituma, Favour Ngozi Okoro (2026). Impact of Data Preprocessing and Missing-Value Handling on Machine Learning-Based Heart Disease Prediction. Cosmo Research & Science International Journal, Jul-25(1), 607-627.
Rebecca Nneka Oguzor, Dr. Chinagolum Ituma, Favour Ngozi Okoro. "Impact of Data Preprocessing and Missing-Value Handling on Machine Learning-Based Heart Disease Prediction." Cosmo Research & Science International Journal, vol. Jul-25, no. 1, 2026, pp. 607-627.
@article{CRSIJ26000336,
author = {Rebecca Nneka Oguzor, Dr. Chinagolum Ituma, Favour Ngozi Okoro},
title = {Impact of Data Preprocessing and Missing-Value Handling on Machine Learning-Based Heart Disease Prediction},
journal = {Cosmo Research and Science International Journal},
year = {2025},
volume = {2},
number = {1},
pages = {607-627},
issn = {3108-1584},
url = {https://cosmorsij.com/published/CRSIJ26000336.pdf},
abstract = {The reliability of machine learning-based heart disease prediction depends not only on the choice of predictive models but also on the quality and preparation of the underlying clinical data. Missing values, inconsistent variable representations, and differences in feature scales can affect the reliability of predictive models and their ability to distinguish between patients with and without heart disease. This study investigated the impact of systematic data pre-processing and missing-value handling on machine learning-based heart disease prediction. The study employed a hybrid methodological framework combining the Cross-Industry Standard Process for Data Mining (CRISP-DM) and Object-Oriented Methodology (OOM). CRISP-DM guided the data-driven process from data understanding and preparation to modelling and evaluation, while OOM supported the modular design and implementation of the system components. The study utilized the Cleveland Heart Disease dataset obtained from the University of California Irvine (UCI) Machine Learning Repository. A structured pre-processing pipeline involving data quality assessment, missing-value treatment using K-Nearest Neighbours (KNN) imputation, categorical variable encoding, and feature normalization using Standard Scalar was applied to prepare the dataset for predictive modelling. Logistic Regression, Random Forest, and Boost models were evaluated using accuracy, precision, recall, F1-score, and ROC-AUC. The ROC-AUC results showed that Logistic Regression achieved the highest discriminative performance with an AUC of 0.591, followed by Boost with 0.566, while Random Forest recorded the lowest AUC of 0.514, indicating performance close to random classification. These findings demonstrate that classification accuracy alone may not adequately reflect a model's discriminative capability and highlight the importance of using multiple evaluation metrics in clinical prediction. The study provides a structured and reproducible pre-processing and modelling framework for preparing clinical heart disease data and supports more rigorous evaluation of machine learning approaches for heart disease prediction.},
keywords = {Heart disease prediction, machine learning, data pre-processing, missing-value handling, KNN imputation, CRISP-DM, Object-Oriented Methodology, ROC-AUC, clinical data},
month = {July}
}