Skip to content
Open access

AFFUNet: adaptive feature fusion Transformer U-Net with joint loss function for medical image segmentation

Jul 2026 · Quantitative Imaging in Medicine and Surgery · Vol 16, pp. 697-697 · 0 citations · 31 references
Medicine

TL;DR

This work proposes Adaptive Feature Fusion U-Net (AFFUNet), an adaptive feature fusion (AFF) Transformer U-Net with a joint loss function, aiming to improve segmentation accuracy through dynamic feature fusion, hard sample reweighting, and explicit boundary optimization.

Abstract

Background High-precision medical image segmentation is crucial for enabling computer-aided accurate diagnosis and personalized treatment planning. Although Transformer-based methods have shown promising results, they still face significant challenges. These include a fixed feature fusion strategy that cannot adapt to dynamic changes in feature importance, the standard Dice loss that treats all pixels equally and struggles to focus on difficult samples, and inadequate boundary precision optimization, which limits the clinical applicability of segmentation results. To address these challenges, this study proposes Adaptive Feature Fusion U-Net (AFFUNet), an adaptive feature fusion (AFF) Transformer U-Net with a joint loss function, aiming to improve segmentation accuracy through dynamic feature fusion, hard sample reweighting, and explicit boundary optimization. Methods Based on the Cross-Shaped Window Transformer UNet (CSWin-UNet) architecture, we propose three modifications to address the challenges discussed above: (I) an AFF module that replaces standard skip connection concatenation to dynamically learn fusion weights for encoder and decoder features; (II) focal Dice loss, which adjusts the Dice loss based on prediction confidence to better prioritize difficult samples; (III) boundary-aware loss, which explicitly enhances boundary precision through a gradient-based boundary extraction technique. Results We evaluated our method on two benchmark datasets: the Synapse multi-organ computed tomography (CT) dataset and the Automated Cardiac Diagnosis Challenge (ACDC) cardiac magnetic resonance imaging (MRI) dataset. On the Synapse dataset, our method achieved an average Dice coefficient of 81.21% and an average 95th percentile Hausdorff distance (HD95) of 18.41 mm, representing a 0.09% improvement in Dice and a 0.45 mm reduction in HD95 compared to the state-of-the-art CSWin-UNet. On the ACDC dataset, our method achieved a mean Dice of 89.77%, consistently outperforming CSWin-UNet (88.44%). Ablation studies confirmed that each component contributes to the overall performance, particularly in small organs and boundary regions. Notably, the added AFF module introduces negligible computational overhead: compared to CSWin-UNet, the parameters and floating point operations per second (FLOPs) remained virtually unchanged (−0.02 M, −0.036 G). Thus, our method achieves improved segmentation accuracy while maintaining practical efficiency. Conclusions This work integrates AFF, focal Dice loss, and boundary-aware loss into a collaborative framework that works across the feature, objective, and boundary levels. The resulting end-to-end mechanism achieves improved performance on standard benchmarks while maintaining practical efficiency, warranting further evaluation in clinical settings.

Read PDF

Similar papers

Preprint Aug 2026

CiUNet: A Hybrid Swin-CNN UNet for Medical Image Segmentation

Medical image segmentation requires high accuracy and robustness, yet practical commercial deployment also demands privacy preservation and computational efficiency. In this context, the U-Net architecture, which can be inherently decoupled into independent encoder and decoder components, serves as a natural commercial...

Bin Dong, Jing-Hong Chen · 0 citations
Open access Aug 2026

Attention-Enhanced Bimodal 3D Medical Image Segmentation with Two-Stage Learning

This work proposes an enhanced 3D segmentation framework, UAtten-Unetr, designed to improve segmentation accuracy and robustness in complex medical scenarios, and innovatively developed a unified loss function based on bimodal modality-specific Dice constraints and uncertainty regularization, optimized for synchronous...

Meng-Xuan Li, Hao-Yu Wang · 0 citations
Sep 2026

LMDAU-Net: An Effective Lightweight Multi-scale Deformation Aggregation U-Net for Skin Lesion Segmentation.

An effective Lightweight Multi-scale Deformation Aggregation U-Net (LMDAU-Net), which consists of a Lightweight Local-global Learning Module (LLM) and an Adaptive Interactive Fusion Module (AIF), which utilizes the two branches of the proposed LLM to efficiently learn local fine-grained and global coarse-grained featur...

Jun-Han Hu, Ming Liu, Jing Yang et al. · 0 citations
Open access Aug 2026

TransCat: a hybrid CNN-transformer network with KAN for medical image segmentation

TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.

Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al. · 0 citations
Aug 2026

GLNet: global-to-local aware hybrid framework for medical image segmentation

GLNet adopts a dual-branch encoder that combines a CNN-based Local Detail Perception Branch with a Mamba-based Global Context Modeling Branch, enabling the joint extraction of fine-grained local features and long-range semantic representations.

Dengdi Sun, Long-Long Liu, Xiao-Wei Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.