Skip to content

Author

Hong-Sen Yang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

FSAUnet: Anisotropy-Aware Focal Modulation and Class-Aware MAC-Loss for 3-D Medical Image Segmentation

Accurate three-dimensional medical image segmentation underpins computer-aided diagnosis and surgical planning, yet it is challenged by the spatial anisotropy of clinical computed tomography and magnetic resonance scans and by the topological fragility of fine anatomical structures, such as hepatic vessels, that conventional voxel-overlap supervision may not adequately represent. Conventional dense global self-attention introduces a quadratic voxel-pair interaction term. To address these issues, we propose FSAUnet, a parameter-efficient three-dimensional segmentation framework built upon the self-configuring nnU-Net pipeline and featuring three innovations. First, an adaptive anisotropy-aware focal modulation module aggregates heterogeneous multi-scale context through depth-wise convolutions, including an anisotropic in-plane branch, without constructing a voxel-wise affinity matrix. For fixed batch size, channel width, kernel size, and branch configuration, its theoretical arithmetic cost and principal inference-forward working-memory requirement under sequential branch evaluation scale linearly with the number of spatial voxels. Second, a cascaded shuffle attention module embedded in the skip connections suppresses background noise through a sequential channel-shuffle and spatial-refinement pipeline. Third, a multi-component class-aware loss augments cross-entropy and region-overlap supervision with a boundary-band term and a class-aware clDice term that is selectively activated for tubular structures to encourage topology-aware segmentation. The class-specific coefficients are fixed rather than learned or sample-adaptive, and a lower-bounded denominator is used only as a numerical safeguard when the soft-skeleton volume becomes small. Under full five-fold internal cross-validation against nnU-Net on three computed tomography datasets from the Medical Segmentation Decathlon, FSAUnet achieves mean Dice scores of 66.83, 70.16, and 53.84 percent on Task07, Task08, and Task10, respectively, corresponding to descriptive mean differences of 1.52, 1.32, and 4.50 percentage points. No paired case-level significance test was performed. Additional comparisons with representative segmentation architectures and component ablations are conducted using a controlled fixed fold-0 protocol and are interpreted descriptively. FSAUnet requires approximately 22.0 percent fewer parameters than nnU-Net. Controlled module-level profiling is consistent with the analytical spatial-scaling trends of A2FM and explicit dense global self-attention, while whole-network profiling shows that the parameter reduction does not necessarily translate into lower one-patch latency or peak allocated inference memory. Ablation results indicate dataset- and metric-dependent behavior of the evaluated configurations. Overall, FSAUnet obtains higher cross-validation mean Dice with fewer parameters under the evaluated settings, without establishing statistical significance, strict architecture-only superiority, or end-to-end runtime superiority.

Hong-Sen Yang, Ling-Fei Cheng, Yu-Xiang Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.