Skip to content
Open access

Modified DeepLabV3+ architecture with global context integration and supervised contrastive learning for semantic segmentation of ultra-high-resolution image

Sep 2026 · Informatics · 0 citations · 9 references

Abstract

Objectives. We propose a modern method for semantic segmentation of ultra-high-resolution (4K) video frames in the field of Earth remote sensing using a modification of the DeepLabV3+ convolutional neural network. Methods. The method is aimed at solving two critical problems: the limited receptive field of the model when working with local image fragments due to GPU memory constraints, and the poor separation of classes with similar color characteristics and texture. The baseline architecture was supplemented with two attention mechanism modules to improve object localization, and the standard cross-entropy loss function was supplemented with a Supervised Contrastive Learning (SupCon) algorithm for better separation of similar classes. Feature-wise Linear Modulation (FiLM) was embedded into the Atrous Spatial Pyramid Pooling (ASPP) module to introduce global context in the form of features extracted from the original image. Results. The results of the experiments showed that the proposed approach successfully minimizes artifacts at class boundaries, has a high ability to distinguish visually similar classes, and improves segmentation accuracy under the conditions of a specific dataset. The proposed architecture outperforms the baseline by 0,45 % in terms of mean Intersection over Union. Conclusion. The developed modification of DeepLabV3+ effectively addresses the challenges of semantic segmentation of ultra-high-resolution images in the field of Earth remote sensing, achieving high accuracy with a negligible increase in computational complexity.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.