Skip to content
Research Article Open access CC BY 4.0

Machine Learning for Hate Text Speech Detection: A Comprehensive Review of Techniques, Dataset and Challenges

Usman Idris Ismail, Suleiman Salihu Jauro, Nuhu Abdulalim Muhammad, Saadatu Ali Jijji, Joshua C Shawulu, Abdullahi Adam Galadima

Asian Journal of Research in Computer Science · pp. 204–218 · Published 10 Mar 2026

10.9734/ajrcos/2026/v19i2832

Abstract

Conventional moderation practices, which rely on human reviewers to identify and remove harmful contents are often labor-intensive, subjective, and unable to cope with the massive volume of user-generated data produced daily. Hate text speech has become a pervasive challenge across digital platforms, prompting extensive research into automated detection methods capable of identifying harmful and abusive content at scale. This review provides a comprehensive synthesis of machine learning approaches for hate speech detection, examining the linguistic characteristics of hateful expressions, the evolution of datasets, and the progression of modelling techniques from traditional machine learning to deep learning and transformer-based architectures. The analysis highlights the complexity of hate speech as a sociolinguistic phenomenon, particularly in its implicit, coded, and context dependent forms, which remain difficult for automated systems to detect reliably. Significant limitations in existing datasets including annotation inconsistency, class imbalance, domain specificity, and limited multilingual coverage further constrain model performance and generalization. Across the literature, challenges related to bias and inadequate evaluation practices persist. By synthesizing current trends and identifying gaps. This review outlines key research directions focused on contextual modelling, multilingual and cross-cultural resources, implicit hate detection, fairness aware algorithms, and adaptive learning strategies. The findings underscore the need for interdisciplinary with ethically grounded approaches to develop robust and socially responsible hate speech detection systems capable of supporting safer online environments. Overall, hate speech detection remains an evolving field that requires ongoing refinement of datasets, models, and evaluation practices. By addressing current gaps and embracing innovative approaches, future systems can better support the creation of safer and more inclusive digital environments.

Hate speech machine learning dataset annotation implicit hate speech abusive language

Cited by 0

No indexed citations yet.

Article metrics

Real usage data collected on this platform.

0

Page views

0

PDF downloads

0

Outbound clicks

0

Citations

Views by country

Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".

No views recorded yet.

Traffic sources

Referring site, by host.

No traffic recorded yet.

Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.