Performance Comparison of Imputation Methods for Mixed Data Missing at Random with Small and Large Sample Data Set with Different Variability
Kyei Baffour Afari, Christina Nicole Holder Lewis
Asian Journal of Probability and Statistics · pp. 16–39 · Published 1 Oct 2022
10.9734/ajpas/2022/v20i2416Abstract
One of the concerns in the field of statistics is the presence of missing data, which leads to bias in parameter estimation and inaccurate results. However, the multiple imputation procedure is a remedy for handling missing data. This study looked at the best multiple imputation methods used to handle mixed variable datasets with different sample sizes and variability along with different levels of missingness. The study employed the predictive mean matching, classification and regression trees, and the random forest imputation methods. For each dataset, the multiple regression parameter estimates for the complete datasets were compared to the multiple regression parameter estimates found with the imputed dataset. The results showed that the random forest imputation method was the best for mostly a sample of 500 irrespective of the variability. The classification and regression tree imputation methods worked best mostly on sample of 30 irrespective of the variability.
Cited by 1
Rifa Khoirunisa, Ahmad Faisal Sani, Darmawan Lahru Riatma · Brilliance: Research of Artificial Intelligence · 2025
Related research
- Appraising the Impact of Naraj Barrage on Sedimentation of Chilika Lagoon; the Soft Computing Model for Prediction — shares topic coverage
- Assessment of the Different Machine Learning Models for Prediction of Cluster Bean (Cyamopsis tetragonoloba L. Taub.) Yield — shares topic coverage
- A Multi-Dimensional Evaluation Framework for IoT Intrusion Detection: Balancing Accuracy, Efficiency and Real-World Deployment Constraints — shares topic coverage
- Predictive Modeling of Digital Credit Risk in Commercial Banks Using Machine Learning Algorithms — shares topic coverage
- An Overview of the Application of Machine Learning and Deep Learning Techniques for Agricultural Crop Yield Prediction in Terms of Methods, Data Inputs and Prospects — shares topic coverage
Article metrics
Real usage data collected on this platform.
0
Page views
0
PDF downloads
0
Outbound clicks
1
Citations
Views by country
Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".
No views recorded yet.
Traffic sources
Referring site, by host.
No traffic recorded yet.
Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.