Object Detection Algorithms Based on Deep Learning: A Review
Jintao Meng, Shaokai Shen, Jiaqi Wang, Chunjian Zhou
Asian Journal of Research in Computer Science · pp. 1–12 · Published 8 Jul 2024
10.9734/ajrcos/2024/v17i7485Abstract
With the continuous development of deep learning, object detection algorithms based on deep learning have made significant progress in the field of computer vision, widely applied in areas such as autonomous driving, industrial inspection, agriculture, transportation, and medicine. Traditional object detection algorithms face issues such as low detection efficiency and poor robustness. However, deep learning-based object detection algorithms significantly enhance detection accuracy and generalization by learning low-level and high-level image features. This article first introduces traditional object detection algorithms and their existing problems, then elaborates on the main processes, innovations, advantages, disadvantages, and experimental results on datasets of deep learning-based object detection algorithms. It focuses on the development of Two-Stage and One-Stage object detection algorithms, and provides an outlook on the future development of object detection algorithms, discussing challenges such as the coordination of detection speed and accuracy, difficulties in detecting small objects, real-time detection tasks, and multi-modal fusion applications, and proposes possible future directions.
References (53)
- 1 Li Meian. Research on object detection algorithm based on deep learning. Journal of Physics: Conference Series. Vol. 1995. No. 1. IOP Publishing; 2021. [DOI]
- 2 Wang Huijuan. A review of 3D object detection based on autonomous driving.The visual computer. 2024;1-19. [DOI]
- 3 Zhang Haigang. Adaptive visual detection of industrial product defects. Peer J Computer Science. 2023;9:e1264. [DOI]
- 4 Ariza-Sentís, Mar. Object detection and tracking in precision farming: A systematic review. Computers and Electronics in Agriculture. 2024;219:108757. [DOI]
- 5 He, Shouhui. Automatic recognition of traffic signs based on visual inspection. IEEE Access. 2021;9:43253-43261. [DOI]
- 6 Kaur Amrita. A survey on deep learning approaches to medical images and a systematic look up into real-time object detection. Archives of Computational Methods in Engineering. 2021;1-41.
- 7 Lowe, David G. Distinctive image features from scale-invariant keypoints. International journal of computer vision 2004;60:91-110. [DOI]
- 8 Viola Paul, Michael Jones. Rapid object detection using a boosted cascade of simple features. Proceedings of the 2001 IEEE computer society conference on computer vision and pattern recognition. CVPR 2001. Vol. 1. Ieee; 2001. [DOI]
- 9 Dalal Navneet, Bill Triggs. Histograms of oriented gradients for human detection." IEEE computer society conference on computer vision and pattern recognition (CVPR'05). Ieee. 2005;1. [DOI]
- 10 Felzenszwalb Pedro F. Object detection with discriminatively trained part-based models. IEEE transactions on pattern analysis and machine intelligence 2009; 32(9):1627-1645. [DOI]
- 11 Uijlings, Jasper RR. Selective search for object recognition. International Journal of Computer Vision. 2013;104: 154-171. [DOI]
- 12 Girshick Ross. Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE conference on computer vision and pattern recognition; 2014. [DOI]
- 13 Kaiming He. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE transactions on pattern analysis and machine intelligence. 2015; 37(9):1904-1916. [DOI]
- 14 Girshick Ross. Fast r-cnn.Proceedings of the IEEE International Conference on Computer Vision; 2015. [DOI]
- 15 Ren, Shaoqing. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems. 2015;28.
- 16 Dai, Jifeng. R-fcn: Object detection via region-based fully convolutional networks. Advances in Neural Information Processing Systems. 2016;29.
- 17 Kaiming He. Mask r-cnn. Proceedings of the IEEE international Conference on Computer Vision; 2017.
- 18 Cai, Zhaowei, Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2018. [DOI]
- 19 Ouyang, Wanli. Chained cascade network for object detection. Proceedings of the IEEE International Conference on Computer Vision; 2017. [DOI]
- 20 Chen, Kai, et al. "Hybrid task cascade for instance segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2019. [DOI]
- 21 Sermanet, Pierre. Overfeat: Integrated recognition, localization and detection using convolutional networks. Arxiv preprint arxiv: 2013;1312.6229.
- 22 Redmon Joseph. You only look once: Unified, real-time object detection. Proceedings of the IEEE conference on computer vision and pattern recognition; 2016. [DOI]
- 23 Szegedy Christian. Going deeper with convolutions." Proceedings of the IEEE conference on computer vision and pattern recognition; 2015. [DOI]
- 24 Maas, Andrew L, Awni Y, Hannun, Andrew Y. Ng. Rectifier nonlinearities improve neural network acoustic models. Proc. icml. 2013;30(1).
- 25 Redmon Joseph, Ali Farhadi. YOLO9000: better, faster, stronger. Proceedings o the IEEE Conference on Computer Vision and Pattern Recognition; 2017. [DOI]
- 26 Simonyan Karen, Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. Arxiv preprint arxiv. 2014;1409:1556.
- 27 Lin, Min, Qiang Chen, Shuicheng Yan. Network in network. Arxiv preprint arxiv. 2013;1312:4400.
- 28 Redmon, Joseph, Ali Farhadi. Yolov3: An incremental improvement. Arxiv preprint arxiv. 2018;1804:02767.
- 29 Bochkovskiy, Alexey, Chien-Yao Wang, and Hong-Yuan Mark Liao. Yolov4: Optimal speed and accuracy of object detection. arxiv preprint arxiv. 2020; 2004:10934.
- 30 He Kaiming. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE transactions on pattern analysis and machine intelligence 2015; 37(9):1904-1916. [DOI]
- 31 Liu, Shu. Path aggregation network for instance segmentation. Proceedings of the IEEE conference on computer vision and pattern recognition; 2018. [DOI]
- 32 Zheng, Zhaohui. Distance-IoU loss: Faster and better learning for bounding box regression. Proceedings of the AAAI Conference on Artificial Intelligence. 2020;34(7). [DOI]
- 33 Ge Zheng. Yolox: Exceeding yolo series in 2021. Arxiv preprint arxiv. 2021;(2107): 08430.
- 34 Law, Hei, Jia Deng. Cornernet: Detecting objects as paired keypoints. Proceedings of the European conference on computer vision (ECCV); 2018. [DOI]
- 35 Duan, Kaiwen. Centernet: Keypoint triplets for object detection. Proceedings of the IEEE/CVF international conference on computer vision; 2019. [DOI]
- 36 Zhang, Hongyi. Mixup: Beyond empirical risk minimization. arxiv preprint arxiv. 2017;1710:09412.
- 37 Zheng Ge. Ota: Optimal transport assignment for object detection. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition; 2021.
- 38 Chuyi Li, YOLOv6: A single-stage object detection framework for industrial applications. arxiv preprint arxiv: 2022;2209:02976.
- 39 Ding, Aohan. Repvgg: Making vgg-style convnets great again." Proceedings of the IEEE/CVF conference on computer vision and pattern recognition; 2021. [DOI]
- 40 Gevorgyan, Zhora. SIoU loss: More powerful learning for bounding box regression. Arxiv preprint arxiv: 2022; 2205:12740.
- 41 Wang, Chien-Yao, Alexey Bochkovskiy, Hong-Yuan Mark Liao. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition; 2023. [DOI]
- 42 Wang, Chien-Yao, I-Hau Yeh, Hong-Yuan Mark Liao. YOLOv9: Learning what you want to learn using programmable gradient information. Arxiv preprint arxiv: 2024; 2402:13616.
- 43 Liu, Wei. Ssd: Single shot multibox detector." Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, Proceedings, Part I 14. Springer International Publishing; 2016. [DOI]
- 44 Jeong, Jisoo, Hyo Park, Nojun Kwak. Enhancement of SSD by concatenating feature maps for object detection. Arxiv preprint arxiv: 2017;1705.09587. [DOI]
- 45 Cheng-Yang Fu. Dssd: Deconvolutional single shot detector. Arxiv preprint arxiv: 2017;1701.06659.
- 46 Li, Zuoxin, Lu Yang, Fuqiang Zhou. FSSD: feature fusion single shot multibox detector. arxiv preprint arxiv: 2017;1712:00960.
- 47 Shen, Zhiqiang. Object detection from scratch with deep supervision. IEEE transactions on pattern analysis and machine intelligence. 2019;42(2):398-412. [DOI]
- 48 Yi, **gru, Pengxiang Wu, Dimitris N. Metaxas. ASSD: Attentive single shot multibox detector. Computer Vision and Image Understanding. 2019;189: 102827. [DOI]
- 49 Howard, Andrew G. Mobilenets: Efficient convolutional neural networks for mobile vision applications. Arxiv preprint arxiv: 2017;1704.04861.
- 50 Howard, Andrew. Searching for mobilenetv3. Proceedings of the IEEE/CVF international conference on computer vision; 2019. [DOI]
- 51 Sandler, Mark. Mobilenetv2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE conference on computer vision and pattern recognition; 2018. [DOI]
- 52 Zhang, **angyu. Shufflenet: An extremely efficient convolutional neural network for mobile devices. Proceedings of the IEEE conference on computer vision and pattern recognition; 2018. [DOI]
- 53 Ningning Ma. Shufflenet v2: Practical guidelines for efficient cnn architecture design." Proceedings of the European conference on computer vision (ECCV); 2018.
Cited by 6
Boonchoo Jitnupong, Jirawan Charoensuk, Supaporn Bundasak · Proceedings of Engineering and Technology Innovation · 2026
Chengming Song, Lingzhi Wu, B. Gao · IEEE Transactions on Smart Grid · 2025
Yixuan Ma, Yifei Song · 2025 6th International Conference on Intelligent Design (ICID) · 2025
G. Dematti, Uttam Patil, Sushiladevi B. Vantamuri · 2025 Global Conference on Information Technology and Communication Networks (GITCON) · 2025
Aditya Parekh, Maryam Ahmed, Daniel Cachola · Communications in Computer and Information Science · 2025
Zihan Yang, Huiyu Xiang · 2025
Related research
- Automatic Segmentation of Organ at Risk in Head and Neck Cancer CT Images Using Medical Open Network for Artificial Intelligence (MONAI) with Deep Learning Techniques — shares topic coverage
- Diagnostic Accuracy of Artificial Intelligence for Breast Cancer Detection: A Systematic Review — shares topic coverage
- Leveraging Deep Learning Algorithms for Predicting Power Outages and Detecting Faults: A Review — shares topic coverage
- Bridging Prenatal Diagnostics and AI: A Systematic Review and Meta-Analysis of the Efficacy of Advanced Algorithms in Identifying Congenital Fetal Abnormalities — shares topic coverage
- Applications of Deep Learning in Predicting the Risk of Metabolic Syndrome from Lifestyle and Behavioral Factors: A Scoping Review — shares topic coverage
Article metrics
Real usage data collected on this platform.
0
Page views
0
PDF downloads
0
Outbound clicks
6
Citations
Views by country
Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".
No views recorded yet.
Traffic sources
Referring site, by host.
No traffic recorded yet.
Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.