International E-publication: Publish Projects, Dissertation, Theses, Books, Souvenir, Conference Proceeding with ISBN. 

A Quantitative analysis of Machine Learning approaches to missing value imputation

Author Affiliations

  • 1Department of Computer Science, Government First Grade College, Domlur Shanthi Nagar, Bengaluru, India

Res. J. Recent Sci., Volume 15, Issue (3), Pages 68-71, July,2 (2026)

Abstract

Missing data is a persistent challenge in real-world datasets and can significantly reduce the reliability and predictive performance of machine learning models. Accurate imputation strategies are therefore essential for effective data pre processing. This study presents a quantitative evaluation of machine learning–based imputation techniques, with a focused comparative analysis of Naïve Bayes and K-Nearest Neighbor (KNN) methods. The algorithms are assessed using accuracy, root mean square error (RMSE), and computational efficiency across datasets with varying proportions of missing values. Experimental observations indicate that KNN achieves superior estimation accuracy, while Naïve Bayes demonstrates faster execution and better scalability in high-dimensional environments.

References

  1. Lin, W. C., & Tsai, C. F. (2020)., Missing value imputation: A review and analysis of the literature (2006–2017)., Artificial Intelligence Review, 53(2), 1487–1509.
  2. Emmanuel, T., Maupong, T., Mpoeleng, D., Semong, T., Mphago, B., & Tabona, O. (2021)., A survey on missing data in machine learning., Journal of Big data, 8(1), 140.
  3. Hasan, M. K., Alam, M. A., Roy, S., Dutta, A., Jawad, M. T., & Das, S. (2021)., Missing value imputation affects the performance of machine learning: A review and analysis of the literature (2010–2021)., Informatics in Medicine Unlocked, 27, 100799.
  4. Liu, M., Li, S., Yuan, H., Ong, M. E. H., Ning, Y., Xie, F., Saffari, S. E., Shang, Y., Volovici, V., Chakraborty, B., & Liu, N. (2023)., Handling missing values in healthcare data: A systematic review of deep learning-based imputation techniques., Artificial Intelligence in Medicine, 142, 102587.
  5. Rahman, M. G., Islam, M. Z., Islam, M. M., & Uddin, J. (2021)., A systematic review of machine learning-based missing value imputation techniques., Data Technologies and Applications, 55(4), 558–585.
  6. Twala, B. E., Cartwright, M., & Shepperd, M. (2011)., Missing data imputation using the EM algorithm and machine learning techniques., Applied Artificial Intelligence, 25(5), 373–393.
  7. Acuña, E., & Rodriguez, C. (2004)., The treatment of missing values and its effect on classifier accuracy., Classification, Clustering, and Data Mining Applications, 639–647.
  8. García-Laencina, P. J., Sancho-Gómez, J. L., & Figueiras-Vidal, A. R. (2010)., Pattern classification with missing data: A review., Neural Computing and Applications, 19(2), 263–282.
  9. García-Laencina, P. J., Sancho-Gómez, J. L., & Figueiras-Vidal, A. R. (2010)., Missing value imputation on missing completely at random data using multilayer perceptrons., Neural Networks, 24(1), 121–129.
  10. Liu, W., Luo, L., & Zhou, L. (2023)., Online missing value imputation for high-dimensional mixed-type data via generalized factor models., Computational Statistics & Data Analysis, 187, 107822.
  11. Little, R. J. A., & Rubin, D. B. (2019)., Statistical analysis with missing data (3rd ed.)., Wiley.
  12. Rubin, D. B. (1987)., Multiple imputation for nonresponse in surveys., Wiley.
  13. Schafer, J. L. (1997)., Analysis of incomplete multivariate data., Chapman & Hall/CRC.
  14. Van Buuren, S. (2018)., Flexible imputation of missing data (2nd ed.)., Chapman & Hall/CRC.
  15. Stekhoven, D. J., & Bühlmann, P. (2012)., MissForest—Non-parametric missing value imputation for mixed-type data., Bioinformatics, 28(1), 112–118.