ALF-Surg: Atomic language-guided few-shot surgical phase recognition
Abstract
Objective: To enhance the generalization and transferability of visual-language models (VLMs) pre-trained with extensive clinical data for surgical phase recognition (SPR) and workflow analysis tasks, this paper proposes Atomic Language-Guided Few-Shot Surgical Phase Recognition (ALF-Surg), a lightweight framework designed for efficient adaptation of VLMs to diverse procedures and clinical sites under few-shot settings. Methods: We leverage the rich medical knowledge contained in large language models to transform high-level annotations into vision-based atomic action descriptions, thereby injecting rich spatiotemporal cues and instrument-tissue interaction semantics. A fine-grained multimodal fusion module is applied to jointly encode vision–language features, forming stable and discriminative atomic-level and class-level prototypes. Furthermore, we propose a multi-level multimodal matching strategy, which performs video-video and video-text alignment at both atomic and phase levels to ensure robust decision-making. Results: ALF-Surg was extensively evaluated on three multi-institutional and multi-procedure datasets, including Cholec80, BernBypass70, and StrasBypass70. Compared with zero-shot transfer, ALF-Surg achieves significant accuracy improvements using only eight labeled frames per class (1-shot gains of +13.65%, +30.00%, and +35.03%, respectively). Compared with state-of-the-art few-shot baselines, ALF-Surg consistently establishes new performance records across these datasets. Conclusion: Experimental results demonstrate that ALF-Surg achieves robust SPR with minimal supervision across diverse surgical settings.