Performance Comparison of YOLOv8 and YOLOv11 for Manga Text Area Detection Based on the Intersection over Union Metric
Abstract
Manga has complex visual characteristics, such as variations in speech bubble shapes, diverse text orientations, and dense background illustrations, which complicate the automatic text area detection process. Detection errors in the form of false positives and false negatives can cause text areas to be localized inaccurately and affect the processing at subsequent stages. This study compares the performance of YOLOv8m and YOLOv11m in detecting text areas in Japanese manga using the Intersection over Union (IoU) metric. The dataset consists of 551 manga images annotated into three classes, namely clean_text, messy_text, and text_bubble. Both models were trained under the same parameter configuration for 60 epochs to ensure an objective comparison. The evaluation was performed on 50 test images covering 564 text objects. The test results show that YOLOv11m obtained an average IoU of 0.7598, which is higher than YOLOv8m (0.7196). In addition, YOLOv11m exhibited a faster inference time of 1158.72ms compared with 1276.23ms for YOLOv8m. Based on these results, YOLOv11m demonstrated superior performance over YOLOv8m in terms of both localization accuracy and computational efficiency for the Japanese manga text area detection task.
References
Al-Ibrahim, N., Al-Ibrahim, M., & Al-Awadhi, W. (2024). Text Extraction and Recognition in Manga Comics Using Image Processing Techniques. Science and Information Conference, 408–418.
Alif, M. A. R. (2024). Yolov11 for vehicle detection: Advancements, performance, and applications in intelligent transportation systems. ArXiv Preprint ArXiv:2410.22898.
Auliawan, A. G., & Ratna, M. P. (2024). Strategi Cool Japan yang baru untuk menyelamatkan pertumbuhan ekonomi dan citra populer Jepang di era digital. KIRYOKU, 8(2), 590–604.
Del Gobbo, J., & Matuk Herrera, R. (2020). Unconstrained text detection in manga: A new dataset and baseline. ECCV 2020 Workshops.
Ghaffar Nia, N., Kaplanoglu, E., & Nasab, A. (2023). Evaluation of artificial intelligence techniques in disease diagnosis and prediction. Discover Artificial Intelligence, 3(1), 5.
Hidayatullah, P., Syakrani, N., Sholahuddin, M. R., Gelar, T., & Tubagus, R. (2025). Yolov8 to yolo11: A comprehensive architecture in-depth comparative review. ArXiv Preprint ArXiv:2501.13400.
Jaiswal, A., Liu, S., Chen, T., Wang, Z., & others. (2023). The emergence of essential sparsity in large pre-trained models: The weights that matter. Advances in Neural Information Processing Systems, 36, 38887–38901.
Li, Y., Yan, H., Li, D., & Wang, H. (2024). Robust miner detection in challenging underground environments: An improved yolov11 approach. Applied Sciences, 14(24), 11700.
Mahadevkar, S. V, Khemani, B., Patil, S., Kotecha, K., Vora, D. R., Abraham, A., & Gabralla, L. A. (2022). A review on machine learning styles in computer vision—techniques and future directions. Ieee Access, 10, 107293–107329.
Mohammed, M., & Alsunosi, R. (2022). Effect of selecting validation dataset on building random forest and decision tree models. AlQalam Journal of Medical and Applied Sciences, 470–478.
Muclis, P. A., Kharisma, I. L., & others. (2023). Penerapan Algoritma Backpropagation Untuk Text Recognition Yang Ditranslate Ke Bahasa Daerah. Jurnal Informatika Dan Rekayasa Perangkat Lunak, 5(1), 21–32.
Puteri, N. R., & Meirza, A. (2024). Implementasi Metode YOLOV5 dan Tesseract OCR untuk Deteksi Plat Nomor Kendaraan. Journal of Computer Science and Visual Communication Design, 9(1), 424–435.
Reza, R. A., Bintoro, A., Multazam, T., Hasibuan, A., & Badriana, B. (2025). Analysis comparison effect of image capture speed on rice pest detection using yolov 5 and yolov 7. Journal Geuthee of Engineering and Energy, 4(2), 81–91.
Shah, R. (2024). Global GPU Network (GGN): Performance Benchmarking and Comparative Analysis Against Traditional GPU Computing Models.
Shia, W.-C., & Ku, T.-H. (2024). Enhancing microcalcification detection in mammography with YOLO-v8 performance and clinical implications. Diagnostics, 14(24), 2875.
Silmina, E. P., Sunardi, S., Yudhana, A., & others. (2025). Comparative Analysis of YOLO Deep Learning Model for Image-Based Beef Freshness Detection. JITK (Jurnal Ilmu Pengetahuan Dan Teknologi Komputer), 11(1), 250–265.
Sirisha, U., Praveen, S. P., Srinivasu, P. N., Barsocchi, P., & Bhoi, A. K. (2023). Statistical analysis of design aspects of various YOLO-based deep learning models for object detection. International Journal of Computational Intelligence Systems, 16(1), 126.
Stodt, J., Reich, C., & Clarke, N. (2023). Unified intersection over union for explainable artificial intelligence. Proceedings of SAI Intelligent Systems Conference, 758–770.
Tjahyanti, L. P. A. S., & Pratama, P. A. (2025). Deteksi Objek Menggunakan Convolutional Neural Network (CNN) pada Pengolahan Citra Digital. KOMTEKS, 4(2), 35–40.
Wu, S., Yang, J., Wang, X., & Li, X. (2022). Iou-balanced loss functions for single-stage object detection. Pattern Recognition Letters, 156, 96–103.
Zaidi, S. S. A., Ansari, M. S., Aslam, A., Kanwal, N., Asghar, M., & Lee, B. (2022). A survey of modern deep learning based object detection models. Digital Signal Processing, 126, 103514.
Zhang, H., Xu, C., & Zhang, S. (2023). Inner-IoU: more effective intersection over union loss with auxiliary bounding box. ArXiv Preprint ArXiv:2311.02877.
Zou, Z., Chen, K., Shi, Z., Guo, Y., & Ye, J. (2023). Object Detection in 20 Years: A Survey. Proceedings of the IEEE, 111(3), 257–276.





