Skip to content
Open access

MonuSegFormer: A Hybrid Swin-Transformer Architecture for Semantic Segmentation of Moroccan Cultural Heritage Monuments

Sep 2026 · Journal of Imaging · 0 citations · 48 references

Abstract

Semantic segmentation of architectural elements in cultural heritage sites lies at the intersection of computer vision and digital preservation. Moroccan historical monuments spanning mosques, madrasas, royal gates (babs), and mausoleums across Fez, Rabat, Marrakech, Meknes, and Tetouan present unique challenges, including extreme texture ambiguity between weathered wall surfaces and background, pronounced multi-scale variation from individual window and door openings to full rooftop surfaces spanning tens of metres, and severe class imbalance. We introduce MonuSegFormer, a heritage-specific hybrid architecture coupling a pretrained Swin-B encoder (ImageNet-22K) with a Multi-Scale Atrous Fusion (MSAF) module and a CBAM-augmented progressive decoder. Evaluated on the Moroccan Monuments Dataset (MMD, 2686 annotated RGB images, 5 semantic classes, 5 cities) on the held-out test set (403 images), MonuSegFormer achieves mIoU = 81.7%, outperforming SegFormer-B5 (76.4%), Mask2Former (78.9%), and DeepLabV3+ (71.3%). A composite CE+Dice+Focal loss with class-balanced weights addresses severe class imbalance, yielding the largest gains on the minority classes Door (+5.3 p.p.) and Window (+5.3 p.p.) over the next-best Mask2Former. All qualitative results are produced by real trained-model inference on held-out test images.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.