Ensemble Machine Learning for Student Performance Prediction in Higher Education: A Critical Review of Predictive Gains, Methodological Quality and Deployment Constraints
Okwedi Kelicha, Ugochukwu Febechi Blessing, Adanna Anyanwu, Chukwueke Nwagbara
Asian Journal of Research in Computer Science · pp. 133–153 · Published 17 Sep 2026
10.9734/ajrcos/2026/v19i9911Abstract
Ensemble machine learning has become the default modelling strategy in research on the prediction of student academic performance in higher education, and published comparisons routinely report that bagged, boosted and stacked combinations of base classifiers outperform single learners. The accumulated literature nevertheless remains difficult to interpret, because reported gains derive from heterogeneous institutional datasets, inconsistent evaluation protocols and outcome definitions that range from continuous grade estimation to binary at-risk flagging. This review critically appraises the evidence on ensemble approaches to student performance prediction in tertiary settings, with attention to the conditions under which ensembling delivers genuine improvement rather than apparent improvement generated by evaluation design. Peer-reviewed literature was identified through structured searching of five scholarly databases and indexes, supplemented by backward and forward citation searching, with selection based on relevance, methodological adequacy and contribution to the review question rather than citation volume alone. Five themes structure the synthesis: the mechanisms by which ensembles reduce predictive error in educational data; the construction of predictor sets and the trade-off between earliness and accuracy; the methodological adequacy of the evidence base, including resampling practice, metric selection and leakage risk; the transferability of models across courses, cohorts and institutions; and the translation of predictions into interpretable, equitable and actionable institutional practice. The evidence supports a modest and context-dependent ensemble advantage over well-tuned single learners, most reliably for heterogeneous tabular predictors and moderate sample sizes, and least reliably where evaluation protocols are weak. Reported accuracies in the upper ninetieth percentile are frequently associated with design features that inflate apparent performance. Confidence in the field's central claims is limited by single-institution samples, inconsistent reporting and a scarcity of evidence that predictions change student outcomes. Priorities include multi-institution external validation, prospective evaluation of intervention effects, standardised reporting of temporal validation, and fairness auditing across the full deployment cycle.
Cited by 0
No indexed citations yet.
Related research
- Forecasting the Risk Factors of COVID-19 through AI Feature Fitting Learning Process (FfitL-CoV19) — shares topic coverage
- Application of Stacking-Based Ensemble Learning Model for Water Quality Prediction — shares topic coverage
- Prediction Error Reduction Function as a Variable Importance Score — shares topic coverage
- Ensemble Learning Techniques for Rice Nutrient Disease Deficiency Detection and Prediction Analysis — shares topic coverage
- A Novel and Effective Multi-Model-Based Default Risk Analysis and Prediction in the Business Sector — shares topic coverage
Article metrics
Real usage data collected on this platform.
0
Page views
0
PDF downloads
0
Outbound clicks
0
Citations
Views by country
Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".
No views recorded yet.
Traffic sources
Referring site, by host.
No traffic recorded yet.
Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.