Deep Neural Network-Based Image Segmentation: Architectural Comparison, Optimization, and Performance Evaluation
Keywords:
Autoencoder, Computer Vision, Deep Learning, Image Segmentation, Neural Networks, U-Net ArchitectureAbstract
Image segmentation is a key step in many computer vision systems, particularly in medical imaging, remote sensing, robotics, and precision agriculture. Over the years, a wide range of traditional approaches have been developed for image segmentation, such as binary thresholding Otsu's method, the watershed method, and K-means clustering, which leverage domain-specific knowledge to efficiently solve segmentation problems in particular application contexts. Unfortunately, however, these methods show significant limitations when faced with the increasing complexity of visual scenes and data heterogeneity. In reality, the fundamental principle of image segmentation is to classify each pixel based on similar attributes. The goal of this pixel-by-pixel classification is to achieve a process that is global, relatively insensitive to noise, locally accurate, and incorporates multi-scale filtering mechanisms. However, traditional methods cannot simultaneously meet these requirements. Faced with the limitations of traditional methods, the emergence of deep learning approaches, particularly those based on neural networks and especially convolutional neural networks, has established itself as an effective solution for image segmentation. The objective of this study is to comparatively evaluate and optimize different deep learning architectures to improve the accuracy and robustness of image segmentation in complex visual environments. The study evaluates the performance of models such as the linear network, Simple U-Net, Extended U-Net, U-Net, DeepLabV3+, and SegFormer, incorporating optimization strategies including data augmentation, fine-tuning, the addition of attentional mechanisms, and the proposal of a specific metric to track performance changes during neural network model training. The experimental results demonstrate improved segmentation performance across the evaluated architectures, as measured by Dice, Intersection over Union (IoU), Precision, and Recall. Among the evaluated models, DeepLabV3+ achieved the highest Dice score (0.91), IoU (0.84), and Precision (0.92), while SegFormer achieved the highest Recall (0.91). These findings demonstrate the effectiveness of deep neural network architectures for accurate pixel-level segmentation and provide a comparative basis for selecting suitable architectures for complex visual environments.
Downloads
References
[1] Long, J., Shelhamer, E., & Darrell, T. (2017). Fully convolutional networks for semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(4), 640--651. https://doi.org/10.1109/TPAMI.2016.2572683.
[2] Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention: MICCAI 2015 (pp. 234--241). Springer. https://doi.org/10.1007/978-3-319-24574-4_28.
[3] Garcia-Garcia, A., Orts-Escolano, S., Oprea, S. O., Villena-Martinez, V., & Garcia-Rodriguez, J. (2018). A review on deep learning techniques applied to semantic segmentation. arXiv. https://doi.org/10.48550/arXiv.1704.06857.
[4] Minaee, S., Boykov, Y. Y., Porikli, F., Plaza, A. J., Kehtarnavaz, N., & Terzopoulos, D. (2021). Image segmentation using deep learning: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(7), 3523--3542. https://doi.org/10.1109/TPAMI.2021.3059968.
[5] Pandey, J. K., Verma, S. K., Kumar, J., Perwej, Y., Jha, S. K., Panchal, B. Y., Ferdouse, R., Sindhu, V., Banerjee, S., Tiwari, M., Badal, R., Mandal, P., & Baghel, J. S. (2026). Transformative role of advanced neural computation in clinical image diagnostics: A review of key concepts and applications. *Seminars in Ultrasound, CT and MRI*. Advance online publication. https://doi.org/10.1053/j.sult.2026.06.010
[6] Otsu, N. (1979). A threshold selection method from gray-level histograms. IEEE Transactions on Systems, Man, and Cybernetics, 9(1), 62--66. https://doi.org/10.1109/TSMC.1979.4310076.
[7] Vincent, L., & Soille, P. (1991). Watersheds in digital spaces: An efficient algorithm based on immersion simulations. IEEE Transactions on Pattern Analysis and Machine Intelligence, 13(6), 583--598. https://doi.org/10.1109/34.87344.
[8] Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018). Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 801--818). Springer. https://doi.org/10.1007/978-3-030-01234-2_49.
[9] Zhao, H., Shi, J., Qi, X., Wang, X., & Jia, J. (2017). Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2881--2890). IEEE. https://doi.org/10.1109/CVPR.2017.660.
[10] Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J. M., & Luo, P. (2021). SegFormer: Simple and efficient design for semantic segmentation with transformers. In Advances in neural information processing systems (Vol. 34, pp. 12077--12090).
[11] LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521, 436--444. https://doi.org/10.1038/nature14539.
[12] Baron, D. (2019). Machine learning in astronomy: A practical overview. Annual Review of Astronomy and Astrophysics, 57, 1--34.
[13] https://doi.org/10.1146/annurev-astro-081817-051840.
[14] Girshick, R., Donahue, J., Darrell, T., & Malik, J. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 580--587). IEEE. https://doi.org/10.1109/CVPR.2014.81
[15] Gonzalez, R. C., & Woods, R. E. (2018). Digital image processing (4th ed.). Pearson.
[16] MacQueen, J. (1967). Some methods for classification and analysis of multivariate observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability (Vol. 1, pp. 281--297). University of California Press.
[17] LeCun, Y., & Bengio, Y. (1998). Convolutional networks for images, speech, and time series. In M. A. Arbib (Ed.), The handbook of brain theory and neural networks (2nd ed., pp. 255--258). MIT Press.
[18] LeCun, Y., Bottou, L., Orr, G. B., & Müller, K.-R. (1998). Efficient backprop. In G. B. Orr & K. Müller (Eds.), Neural networks: Tricks of the trade (pp. 9--50). Springer. https://doi.org/10.1007/3-540-49430-8_2
[19] Cireşan, D. C., Giusti, A., Gambardella, L. M., & Schmidhuber, J. (2012). Deep neural networks segment neuronal membranes in electron microscopy images. In Advances in neural information processing systems (Vol. 25, pp. 2843--2851).
[20] Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, M., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment Anything. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 4015--4026.
[21] Li, X., Ding, H., Yuan, H., Zhang, W., Pang, J., Cheng, G., Chen, K., Liu, Z., & Loy, C. C. (2024). Transformer-Based Visual Segmentation: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12), 10138--10163. https://doi.org/10.1109/TPAMI.2024.3434373
[22] Zhou, T., Xia, W., Zhang, F., Chang, B., Wang, W., Yuan, Y., Konukoglu, E., & Cremers, D. (2024). Image Segmentation in Foundation Model Era: A Survey. arXiv preprint arXiv:2408.12957.
[23] Zhang, Y., Shen, Z., & Jiao, R. (2024). Segment Anything Model for Medical Image Segmentation: Current Applications and Future Directions. Computers in Biology and Medicine, 171, 108238. https://doi.org/10.1016/j.compbiomed.2024.108238
[24] Lee, H. H., Gu, Y., Zhao, T., Xu, Y., Yang, J., Usuyama, N., Wong, C., Wei, M., Landman, B. A., Huo, Y., Santamaria-Pang, A., & Poon, H. (2024). Foundation Models for Biomedical Image Segmentation: A Survey. arXiv preprint arXiv:2401.07654.
[25] Zhou, T., et al. (2024). Image Segmentation in Foundation Model Era: A Survey. arXiv preprint arXiv:2408.12957.
[26] Pandey, J. K., Kumar, S., Lamin, M., Gupta, S., Dubey, R. K., & Sammy, F. (2022). A metaheuristic autoencoder deep learning model for intrusion detector system. *Mathematical Problems in Engineering*, *2022*, Article 3859155. https://doi.org/10.1155/2022/3859155
[27] From CNN to Transformer: A Review of Medical Image Segmentation Models (2024). This recent review compares representative CNN- and Transformer-based segmentation architectures, including FCN, U-Net, DeepLab, TransUNet, Swin-Unet, and SAM.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal Press

This work is licensed under a Creative Commons Attribution 4.0 International License.
All articles are distributed under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0): anyone may copy, redistribute, remix, transform, and build upon the material in any medium or format for any purpose, including commercially, provided appropriate credit is given to the original authors and a link to the licence is provided.