Fine-grained visual classification (FGVC) plays a crucial role in the realm of computer vision. Recently, multimodal FGVC methods, leveraging textual descriptions as semantic guidance, have gained considerable attention. However, current approaches often encounter two primary limitations: 1) Redundant or ambiguous textual descriptions: existing methods rely on raw or generated descriptions without filtering, introducing redundant and ambiguous semantic noise; and 2) Underutilization of hierarchical visual features: most approaches align single-layer visual features with auxiliary semantic embeddings, underutilizing hierarchical information. To address these challenges, we propose a task-oriented multimodal FGVC framework that eliminates textual redundancy while enhancing multi-layer alignment between cross-modalities. Specifically, our method comprises two key components: Hierarchical Semantic Purification (HSP) and Multi-layer Cross-Modal Alignment (MCA). The former employs a semantic distillation dictionary to eliminate redundant elements and uses a self-attention mechanism for ranking and semantic refinement. The latter establishes effective cross-modal fusion by integrating multi-layer features with purified text features, effectively combining multi-scale visual representations. Experimental results on 7 public datasets demonstrate that our proposed method outperforms existing counterparts, contributing to advancements in fine-grained visual classification.
Meng-Huan Zhang, Qing Cai, Fan Zhang et al.· IEEE Transactions on Image P...· 0 citations
Four-dimensional magnetic resonance fingerprinting (4DMRF) provides multi-parametric and motion-resolved tissue property quantification, promising to enhance the precision of liver cancer radiotherapy. However, its clinical translation is hindered by prolonged reconstruction time. Deep learning acceleration is fundamentally constrained by the lack of ground-truth 4D data. To this end, we propose SS-4DMRF, the first self-supervised reconstruction framework for 4DMRF, to reconstruct motion-resolved tissue maps without using supervised image labels. SS-4DMRF features a core temporal low-rank-constrained registration (TelReg) network for precise motion modeling. It leverages the intrinsic low-rank compressibility of respiratory motion and is directly self-supervised by the original highly undersampled k-space data and the derived subspace images. Motion-resolved tissue maps are reconstructed using a motion-informed compensation approach via a physics-informed pattern matching (PiPM) network. PiPM network incorporates novel multi-scale Swin Transformers with Bloch-equation-guided subspace denoising to achieve high-fidelity tissue quantification. SS-4DMRF was validated on digital phantom (n=30) and in vivo liver cancer patient (n=33) datasets. Compared to state-of-the-art 4DMRF methods, SS-4DMRF demonstrated superior tissue quantification and motion measurement accuracy. It achieved significantly reduced NRMSE in 4D tissue property quantification and improved inter-phase structural repeatability in 4D motion characterization (Paired Student's t-tests, p<0.001). The measured tumor motion trajectory presented strong Peason correlation with motion reference (r=0.939±0.057). Crucially, SS-4DMRF achieves this dual improvement in accuracy with a 10-fold acceleration in reconstruction time compared with conventional 4DMRF methods. By enabling rapid, precise, and motion-resolved quantitative imaging, SS-4DMRF advances the precision of liver cancer radiotherapy and establishes a clinically feasible platform for abdominal quantitative MRI in oncology.
Chen-Yang Liu, Lu Wang, Xiang Wang et al.· Medical Image Anal.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.