Skip to content
Research Article Open access CC BY 4.0

Arabic English Cross-Lingual Plagiarism Detection Based on Keyphrases Extraction, Monolingual and Machine Learning Approach

Mokhtar Al-Suhaiqi, Muneer A. S. Hazaa, Mohammed Albared

Asian Journal of Research in Computer Science · pp. 1–12 · Published 13 Feb 2019

10.9734/ajrcos/2018/v2i330075

Abstract

Due to rapid growth of research articles in various languages, cross-lingual plagiarism detection problem has received increasing interest in recent years. Cross-lingual plagiarism detection is more challenging task than monolingual plagiarism detection. This paper addresses the problem of cross-lingual plagiarism detection (CLPD) by proposing a method that combines keyphrases extraction, monolingual detection methods and machine learning approach. The research methodology used in this study has facilitated to accomplish the objectives in terms of designing, developing, and implementing an efficient Arabic – English cross lingual plagiarism detection. This paper empirically evaluates five different monolingual plagiarism detection methods namely i)N-Grams Similarity, ii)Longest Common Subsequence, iii)Dice Coefficient, iv)Fingerprint based Jaccard Similarity  and v) Fingerprint based Containment Similarity. In addition, three machine learning approaches namely i) naïve Bayes, ii) Support Vector Machine, and iii) linear logistic regression classifiers are used for Arabic-English Cross-language plagiarism detection. Several experiments are conducted to evaluate the performance of the key phrases extraction methods. In addition, Several experiments to investigate the performance of machine learning techniques to find the best method for Arabic-English Cross-language plagiarism detection. According to the experiments of Arabic-English Cross-language plagiarism detection, the highest result was obtained using SVM   classifier with 92% f-measure. In addition, the highest results were obtained by all classifiers are achieved, when most of the monolingual plagiarism detection methods are used. 

Cross language plagiarism detection mono-language plagiarism detection classification machine learning key phrases candidate document

Cited by 8

Siamese GRU for Arabic-English Cross-language Plagiarism Detection

Chaimaa Bouaine, F. Benabbou, Z. Ellaky · 2025 5th International Conference on Innovative Research in Applied Science, Engineering and Technology (IRASET) · 2025

Improving plagiarism detection in text document using hybrid weighted similarity

H. Arabi, M. Akbari · Expert systems with applications · 2022

Cross-Language Plagiarism Detection: Methods, Tools, and Challenges: A Systematic Review

Miguel Botto-Tobar, Alexander Serebrenik, M. van den Brand · International Journal on Advanced Science, Engineering and Information Technology · 2022

Using deep learning models for learning semantic text similarity of Arabic questions

Mahmoud M. Hammad, M. Al-Smadi, Qanita Bani Baker · International Journal of Electrical and Computer Engineering (IJECE) · 2021

Evaluation and Comparison of Cross-lingual Text Processing Pipelines

Robert Jungnickel, André Pomp, Andreas Kirmse · IEEE Symposium Series on Computational Intelligence · 2019

Scalable and language-independent embedding-based approach for plagiarism detection considering obfuscation type: no training phase

Erfaneh Gharavi, H. Veisi, Paolo Rosso · Neural computing & applications (Print) · 2019

A Systematic Review of Multilingual Plagiarism Detection: Approaches and Research Challenges

Chaimaa Bouaine, F. Benabbou, Z. Ellaky · International Journal of Advanced Computer Science and Applications · 2025

Article metrics

Real usage data collected on this platform.

0

Page views

0

PDF downloads

0

Outbound clicks

8

Citations

Views by country

Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".

No views recorded yet.

Traffic sources

Referring site, by host.

No traffic recorded yet.

Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.