Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language models (VLMs) and vision encoders offer an alternative to conventional interaction classifiers by transferring broad visual and semantic knowledge. However, adapting them to fine-grained surgical interactions remains challenging: (1) freezing the vision encoder depends entirely on pretrained representations that may retain noise and provide weak spatial localization, while (2) full fine-tuning can improve global semantic alignment without ensuring that the encoder learns meaningful features in the correct action region. We address these limitations by introducing LAViFiT, an end-to-end latent-action-guided framework for vision-language fine-tuning. An inverse dynamics model captures the visual changes induced by each action, while a forward world model drives the encoder to represent action-relevant regions. A patch-level SIG Regularizer further prevents local feature collapse without additional supervision, such as bounding boxes or pseudo-labels. Experiments across multiple encoders and datasets improve recognition and image-text alignment, while representation analyses show stronger grounding over the complete instrument-tissue interaction region and more spatially coherent features.
Objective: To enhance the generalization and transferability of visual-language models (VLMs) pre-trained with extensive clinical data for surgical phase recognition (SPR) and workflow analysis tasks, this paper proposes Atomic Language-Guided Few-Shot Surgical Phase Recognition (ALF-Surg), a lightweight framework desi...
Hou-Long He, Lin Mao, Cheng-Li Song· Progress in Medical Devices· 0 citations
HCE-CLIP is proposed, a vision–language pretraining framework that formulates medical image–text alignment as a hierarchical label-conditional energy modeling problem and introduces a hierarchical contradiction-based metric that quantifies logical inconsistencies between fine-grained disease predictions and higher-leve...
Cheng-Sheng Mao, Yuan Luo· Proceedings of the Thirty-Fi...· 0 citations
Task-specific lightweight models for surgical phase recognition excel at capturing temporal dynamics but generalize poorly under domain shift. Conversely, surgical foundation models (FMs) offer superior transferability via large-scale pretraining, yet their lack of explicit temporal modeling often yields temporally inc...
Yi-Yi Zhang, Ying Zheng, Wenxin Fan et al.· 1 citation
This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.
Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al.· International Conference on...· 0 citations
Query-based segmentation methods have shown promising potential for surgical instrument segmentation and recognition, which is essential for scene understanding and downstream tasks in computer assisted surgery. However, most existing approaches predominantly rely on per-frame predictions and overlook cross-frame tempo...
Jian Yuan, Wei-Ming Mi, Tao Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.