Aug 2026· Measurement science and technology· Vol 37, pp. 335402· 0 citations· 42 references
Physics
TL;DR
Comprehensive evaluations on multiple datasets and SR scales indicate that the SVRCL-SR achieves superior performance in artifact suppression and high-frequency detail restoration, along with strong robustness.
Abstract
Sparse-view rotational-scanning computed laminography (SVRCL) is an essential technique for the rapid inspection of large-size plate-shaped components. However, its reconstructed results inevitably suffer from severe noise and streak artifacts, which degrade detection reliability. Recently, deep learning-based techniques such as projection domain constraints and image domain post-processing have shown promising application prospects. But two major bottlenecks remain. First, convolutional neural networks struggle to capture the long-range distribution characteristics of streak artifacts. Second, self-attention mechanisms introduce additional computational overhead. To solve the above problems, we devise a dual-domain joint image super-resolution (SR) framework (SVRCL-SR). First, shallow structural features are extracted through convolutional layers. Then, we introduce a feature denoising module that focuses on the mid- to high-frequency regions and dynamically performs dual-domain joint denoising. Finally, large-kernel attention is realized via frequency-domain convolution and element-wise multiplication, which compensates for missing high-frequency information while reducing computational cost. Comprehensive evaluations on multiple datasets and SR scales indicate that the SVRCL-SR achieves superior performance in artifact suppression and high-frequency detail restoration, along with strong robustness.
Deep learning–based multi-contrast Magnetic Resonance (MCMR) super-resolution (SR) has achieved notable success in accelerating image acquisition and improving image quality. However, significant challenges remain when dealing with the volumetric data: 1) Most existing MCMR SR methods primarily rely on single-slice information and fail to exploit high-dimensional volumetric contextual information; 2) Due to the sparsity of the original low-resolution volumetric data, conventional small kernel convolutions struggle to capture long-range contextual information. Although transformer-based approaches can model long-range dependencies, they suffer from high computational and memory demands when applied to high-dimensional volumetric data. To address these challenges, we propose a multi-view large-kernel attention network for MCMR volumetric SR. The method contains three stages: a cross-modality synthesis stage, an inter-slice deformable compensation stage, and a multi-view large-kernel attention fusion stage. Specifically, a multi-view fusion strategy is proposed to exploit the rich spatial contextual information inherent in high-dimensional volumetric data. A large-kernel convolution attention block is proposed to efficiently capture long-range dependencies from the sparsely sampled coronal and sagittal planes. By jointly integrating the high-order multi-view and multi-contrast information, our method successfully reconstructs high-quality MCMR volumetric data. Experimental results across different datasets, along with the downstream segmentation tasks, attest to the effectiveness of the proposed method.
Pengcheng Lei, Juncheng Li, Faming Fang et al.· IEEE Transactions on Computa...· 0 citations
Residual Flow Matching for Image Super-Resolution (RFMSR) is proposed, a vision-only framework that centers the source distribution at the LQ latent, reducing transport distance and preserving structural priors throughout the flow trajectory.
Shuwei Huang, Tianyao Luo, Jicheng Liu et al.· 1 citation
To address the challenge of balancing long-range dependency modeling and detail fidelity in urban scenes, medical slices, and high-resolution remote sensing imagery, this study proposes a lightweight hybrid architecture that integrates a lightweight CNN with a window-based Transformer. In the front-end, depth wise separable convolutions and residual connections are employed to extract edge-sensitive features, complemented by an edge-guided branch and channel recalibration to enhance thin structures and sharp boundaries. In the intermediate stage, multi-scale local window self-attention is utilized to capture long-range dependencies, while learnable windows and sparse global proxies are introduced to reinject global semantic information. In the bridging stage, deformable alignment and gated residual connections are adopted for cross-scale feature fusion, with weights jointly modulated by edge density and global proxies. Experiments are conducted on the Cityscapes validation set under a single-scale inference setting with a resolution of 2048×1024, using an NVIDIA A100 80 GB GPU, batch size of 1, and mixed precision enabled. The proposed method achieves an IoU of 92.3, an F1 score of 93.1, and a speed of 25 FPS, only 12.8 M parameters, 31.6 GFLOPs, and peak GPU memory usage of 2.8 GB. Compared with U-Net (IoU 88.5, F1 89.2, 20 FPS) and SegFormer (IoU 90.1, F1 91.0, 18 FPS), the proposed approach demonstrates clear advantages in accuracy, real-time performance, and intrinsic structural efficiency. The results indicate that the collaboration between global semantics and local details within a unified weighting domain can effectively improve the separability and deploy ability of high-resolution segmentation.
Yuyang Wang, Jiamei Hu, Xinwei Wang et al.· International Conference on...· 0 citations
Processing the lunar astrophotography imagery is challenging since atmospheric turbulence and low-light conditions introduce blur and noise that
obscure small-scale lunar features. The present work proposes an enhancement pipeline that operates entirely within classical signal processing,
providing a transparent alternative to black-box machine learning methods while remaining practical on standard personal computers. The pipeline
integrates four key stages: lucky-imaging frame selection and stacking, wavelet-domain denoising, adaptive local contrast enhancement, and
dual-stage edge sharpening. The pipeline is evaluated on 30 lunar datasets spanning multiple phases and multiple observable conditions. Quantitative
results show Peak Signal-to-Noise Ratio (PSNR) improvements of approximately 6.2–12.4 dB and an increase in Structural Similarity Index
Measure (SSIM) from 0.56 to 0.90, indicating better structural fidelity relative to stacked baselines. Stacking of 16 carefully selected frames yields
effective SNR gains of up to about 3.2 times, while modulation transfer function (MTF) analysis at limb and crater-edge boundaries reveals sharpness
improvements in the order of 28–35%. The workflow requires no specialized accelerators, with typical resource usage of roughly 4.2 MB memory
and 2.3 seconds per megapixel on conventional CPUs. This is demonstrating that classical, interpretable techniques remain highly competitive for
scientific and educational lunar image enhancement.
S. Bhattacharya· Asian journal of applied sci...· 0 citations
To address the challenges of high parameter redundancy and prohibitive computational complexity inherent in traditional convolutional neural networks and Transformer architectures—which impede deployment on resource-constrained edge medical devices—this paper proposes LightVM-SparseUNet, an ultra-lightweight medical image segmentation framework based on state space models. The core innovations are twofold: First, a Multi-path Visual Mamba module is designed to significantly enhance feature extraction efficiency via a linear-complexity inference mechanism while maintaining feature channel integrity. Second, a sparse-sampling self-attention mechanism is integrated into the U-shaped skip connections, enabling the precise capture of long-range spatial dependencies and mitigating spatial information loss at minimal computational cost. Experimental results demonstrate that LightVM-SparseUNet achieves segmentation competitive with state-of-the-art large-scale models across two authoritative public datasets. Critically, the proposed model achieves extreme lightweights, with a parameter count of only 0.08 M and a computational overhead of merely 0.16 GFLOPs.Our method is highly practical, and the code can be found at https://github.com/yjzbkl/LightVM-SparseUNet.
Haojie Fan, Kang Xu, Xiaoyu Hou et al.· Biomedical engineering and p...· 0 citations
Angular super-resolution (ASR) is a fundamental task in light field (LF) imaging, aimed at reconstructing a dense LF from sparsely sampled views. Despite significant progress, current methods often struggle to preserve consistency in complex scenarios such as severe occlusions and large-disparity regions. In this paper, we propose a robust attention-guided multi-dimensional feature fusion network (LFAMF) for LF angular reconstruction. The proposed framework comprises two synergistic stages: a multi-dimensional feature fusion stage and an attention-guided refinement stage. Specifically, we design a multi-stream subnetwork (MFNet) to extract intrinsic physical characteristics across the spatial, angular, EPI, and pseudo-video sequence domains. Simultaneously, a geometry-prior-based subnetwork (GSPNet) is incorporated to leverage scene structure for improved texture preservation. To effectively integrate these complementary streams, an attention-guided fusion subnetwork (AFNet) is employed to adaptively merge intermediate results. Extensive experiments on both synthetic and real-world datasets demonstrate that the LFAMF model significantly outperforms state-of-the-art methods, particularly in maintaining structural integrity at occlusion boundaries and highly textured areas while ensuring superior angular consistency.
Xiyao Hua, B. Su, Daili Yang· PLoS ONE· 0 citations