Skip to content
Review Open access

The Emerging Role of Vision-Language Models in the Automation of Railway Asset Management: A Review and Future Perspective

Jul 2026 · The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences · Vol XLIX-B2-2026, pp. 745-753 · 0 citations · 26 references

TL;DR

This review paper argues that Vision-Language Models (VLMs), a paradigm whose rapid maturation is evidenced by recent comprehensive surveys offer a transformative solution for rail asset management, and provides a focused overview of the limitations of current CV systems.

Abstract

Abstract. The safety, efficiency, and longevity of global railway networks are directly linked to the rigorous inspection and management of their vast inventory of physical assets. Over the past decade, the field has progressed from manual surveys to automated systems leveraging imagery from track-based or aerial platforms. These systems predominantly built on traditional Computer Vision (CV) models have proven effective at detecting a pre-defined set of common assets. However, this progress has exposed a fundamental architectural and operational ceiling: the closed-world assumption. Current models are constrained to a fixed catalogue of classes defined during their training. It makes the model incapable of identifying novel objects or adapting to environmental changes without costly and continuous cycles of data re-annotation, retraining, and redeployment. This review paper argues that Vision-Language Models (VLMs), a paradigm whose rapid maturation is evidenced by recent comprehensive surveys offer a transformative solution. We provide a focused overview of the limitations of current CV systems and map the mechanics of a VLM-powered approach specifically Open-Vocabulary Detection and Reasoning Segmentation directly to the outstanding challenges in rail asset management. Ultimately, the literature suggests that the adoption of VLMs could catalyze a fundamental shift in railway infrastructure management that serves as a key enabler for next-generation Predictive Maintenance and autonomous Digital Twins.

Read PDF

Similar papers

Preprint Aug 2026

A Multi-Sensor Dataset for Monitoring the Operational Environment of Rail Vehicles

Reliable environment monitoring is essential for the safe and efficient operation of automated railway systems, covering all Grades of Automation (GoA), from partially automated (GoA2) to fully automated operation (GoA4). Artificial Intelligence (AI) plays a central role in enabling these systems to detect, classify, and react to potential hazards in real time. The development of such AI-based perception systems requires large volumes of accurately annotated data for training and validation. Within the Digitale Schiene Deutschland (DSD) program, DB InfraGO AG and understandAI GmbH have developed a comprehensive multi- sensor dataset tailored to the needs of railway environment perception. This dataset contains over 7 million high-quality annotations of both railway-specific and general perception objects, captured under varying operational scenarios. The finalized dataset can now be requested at the DB InfraGO AG and serve as a valuable resource for advancing AI-driven environment monitoring in the railway domain.

Claudio Diotallevi, Rodrigo Gudiño, Zaharia Pachalieva et al. · 1 citation
Open access Aug 2026

Development of an RAG-Integrated Agentic BIM System for Intelligent Railway Maintenance

This framework proves an autonomous decision-making system that organically links inspection data with maintenance regulations by transforming static, manual-labor-centered maintenance workflows into intelligent automated models and increases the efficiency of railway infrastructure management while providing a scalable technical foundation for overall asset management of future smart-city infrastructure.

Minjae Jeon, Yong-Gyun Kim, Seok-Han Kim · 0 citations
Review Aug 2026

A GitOps-Driven Annotation Catalog for Fully Automatic Railway Operations

Automatic train operation (ATO) at grade of automation 3 and above (GoA3-GoA4) requires robust AI-based perception systems capable of reliably detecting obstacles and railway-specific objects under real-world conditions. The effectiveness of these modern artificial intelligence approaches depends heavily on large-scale, high-quality, and highly dynamic annotated datasets. However, managing metadata, maintaining provenance, and tracking the iterative evolution of these annotations impose significant infrastructural and regulatory requirements. Existing monolithic data catalogs often suffer from massive operational overhead, poor integration into developer workflows, and severe documentation drift. This paper introduces an innovative, lightweight GitOps-based architecture for metadata management. By leveraging Data-as-Code principles, Continuous Integration/Continuous Deployment (CI/CD) pipelines, and Static Site Generation (SSG), the proposed approach establishes a seamless, developer-centric workflow. This ensures an traceability, enforces strict regulatory compliance, and automatically generates a highly performant dataset overview.

Martin Köppel, Tobias Cronauer, Zekiye Ilknur-Öz et al. · 0 citations
Oct 2026

Visual Foundation Model–Based Multilabel Perception of a Railway Train Operating Environment Using Onboard Surveillance Video

Reliable perception of the train operating environment is essential for supporting efficient and intelligent railway operations. However, traditional trackside sensing infrastructures are sparsely deployed and costly to maintain, making it difficult to obtain whole-process environmental information. To address this challenge, this study proposes a visual foundation model-based multilabel perception framework that leverages existing on-board surveillance videos without requiring additional sensors or manual annotation. The framework utilizes a visual foundation model to generate initial open-vocabulary semantic tags, enabling zero-shot recognition of diverse environmental elements. A semantic refinement mechanism is then introduced to extract controllable environmental labels through similarity matching with a predefined label library. Finally, a dynamic label correction module integrates prior knowledge and temporal cues to suppress frame-level noise and ensure sequence-level consistency. Experiments on a real-world on-board video dataset demonstrate the effectiveness of the proposed framework, achieving an image-level label accuracy of 91.8% and an event-level perception accuracy of 90.6%. This work provides a practical and generalizable pipeline for whole-process railway environment perception and offers new insights into adapting visual foundation models to domain-specific applications.

Shize Huang, Yimin Shen, Qianhui Fan et al. · 0 citations
Open access Aug 2026

An Intelligent Enterprise Asset Management Framework for Railway Maintenance Prioritization: A Machine-learning Proof of Concept Using Benchmark Analogue Data

Railway networks remain capital-intensive and safety-critical, yet maintenance still relies heavily on reactive and calendar-driven regimes that underexploit the condition data modern systems generate. Published machine-learning applications remain fragmented across asset classes and weakly connected to enterprise decision systems, leaving a gap this study addresses by developing an integrated framework uniting enterprise asset management, predictive analytics, and digital analytics for maintenance prioritisation and decision support. Adopting a quantitative design with a design-science orientation, the framework was implemented across stages spanning data preparation, feature engineering, modelling, evaluation, reliability translation, and decision integration. Two established condition sources supplied the binary classification and degradation estimation tasks. Random forest, extreme gradient boosting, support vector machines, and neural networks were trained with feature scaling, with synthetic minority oversampling applied within training folds for the classification task, then assessed through held-out testing and five-fold resampling, stratified for classification and partitioned at engine level for degradation. Gradient boosting achieved the strongest classification balance, recording an F1-score of 0.737 and a ROC-AUC of 0.981, while the neural network minimised degradation error at a grouped five-fold mean RMSE of 47.25 cycles. Resampling confirmed stability, attribution exposed speed, tool wear, and thermal differential as dominant drivers, and all ten highest-ranked observations were confirmed failures, a precision-at-10 of 1.00. Although benchmark analogues constrain generalisation, the contribution is advanced as a proposed architecture supported by a computational proof of concept, in which predictions are converted into auditable, risk-ranked maintenance priorities pending naturalistic validation on operational railway records.

Adeyemi Adebukunola Ishekwene · 0 citations
Jul 2026

Advancing Automatic Recognition through Digital Transformation:Balancing Automation with a Human-Centric Approach

The digital transformation (DT) of academic credential recognition aims to contribute to the policy goal of Automatic Recognition (AR), a long-standing commitment established through normative frameworks like the Lisbon Recognition Convention (LRC), the European Higher Education Area (EHEA) and the Council Recommendation on promoting automatic mutual recognition of qualifications and learning periods abroad. While DT provides key opportunities for efficiency, transparency, and consistency, AR itself - defined primarily as a reduction of separate recognition procedures - must be explicitly differentiated from full technological automotion, a risk arising from over-reliance on digitalization. This article first reconstructs the history of AR policies and subsequently analyzes the binding (inter) national regulations governing DT and Artificial Intelligence (AI), specifically noting that the EU AI Act classifies AI systems used in education and qualification evaluation as high-risk. The research question addresses how DT can successfully support AR’s procedural streamlining without escalating into total automotion that bypasses the principles of human-centric evaluation and the mandatory requirement for human oversight imposed by international frameworks (EU, Council of Europe, UNESCO). Adopting normative document analysis and utilizing the case study of CIMEA (the Italian ENIC-NARIC center), the study identifies crucial pillars for a compliant DT strategy that effectively supports AR: Human Oversight and Accountability, Human-Centric Design, Ethics-by- Design and Quality-by-Design, Robust Data Governance and Privacy, Transparency and Explainability, and AI Literacy/Upskilling for staff.

Luca Ferranti, Ch. Finocchietti, Serena Spitalieri · 0 citations