Visual Foundation Model–Based Multilabel Perception of a Railway Train Operating Environment Using Onboard Surveillance Video
Reliable perception of the train operating environment is essential for supporting efficient and intelligent railway operations. However, traditional trackside sensing infrastructures are sparsely deployed and costly to maintain, making it difficult to obtain whole-process environmental information. To address this challenge, this study proposes a visual foundation model-based multilabel perception framework that leverages existing on-board surveillance videos without requiring additional sensors or manual annotation. The framework utilizes a visual foundation model to generate initial open-vocabulary semantic tags, enabling zero-shot recognition of diverse environmental elements. A semantic refinement mechanism is then introduced to extract controllable environmental labels through similarity matching with a predefined label library. Finally, a dynamic label correction module integrates prior knowledge and temporal cues to suppress frame-level noise and ensure sequence-level consistency. Experiments on a real-world on-board video dataset demonstrate the effectiveness of the proposed framework, achieving an image-level label accuracy of 91.8% and an event-level perception accuracy of 90.6%. This work provides a practical and generalizable pipeline for whole-process railway environment perception and offers new insights into adapting visual foundation models to domain-specific applications.