An Image Segmentation Network Based on Multi-Scale Channels and Optimized Fine-Grained
Keywords:
Fine-grained Segmentation, Computer Vision, Multi-scale Feature Extraction, Small Object Detection, Urban Scene SegmentationAbstract
The rapid advancement of deep learning has greatly propelled the development of semantic segmentation technologies for street scene imagery, which are increasingly critical in applications such as autonomous driving and smart city management. However, when processing high-resolution urban images with complex backgrounds, most existing models struggle to balance the capture of global contextual information with the extraction of local fine-grained details, especially for small or boundary-blurred objects. To overcome these limitations, this paper introduces an enhanced architecture named MSCA-Deeplabv3+, built upon DeepLabV3+ by incorporating a Swin Transformer backbone and a novel Multi-scale Channel Module (MCM). This design significantly improves the model's ability to perceive and segment objects of varying sizes with high precision. Furthermore, to mitigate the issue of sample imbalance—particularly for small objects—a hybrid loss combining cross-entropy and contrastive loss is adopted, enhancing discrimination in challenging regions. Evaluations on the Cityscapes and KITTI datasets demonstrate the effectiveness of our method: on Cityscapes, we achieve a mIoU of 87.3%, and on KITTI, our model reaches 90.3% in accuracy, outperforming several strong baselines. The proposed model also maintains efficient parameter usage with 44.32M parameters, showing strong generalization across different street scenarios.
Published
Issue
Section
License
Copyright (c) 2026 Journal of Information and Computing

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.