Event cameras are increasingly used for Multiple Object Tracking (MOT), but their asynchronous event output often requires specialized methods. Existing processing methods primarily follow two paradigms, pseudo-frames and event-by-event. The former is the prevailing approach since its data format aligns with images, making image-based techniques applicable. However, it suffers from tracking failures when trajectories overlap or are spatially close on pseudo-frames. Facing this challenge, we propose a multi-view pipeline, Multi-view Tracking (MvT), which preserves the 2D data format to leverage image-based techniques directly while introducing additional spatio-temporal views to resolve tracking ambiguities in a single view. MvT comprises a Multi-view Projection (MvP) module and a Multi-view Fusion (MvF) stage. MvP encodes events into three complementary spatio-temporal views while mitigating the pattern discretization. Within MvF, multi-view results are unified into a 3D coordinate system, and tracklets are associated through an optimization model subject to specific criteria combination. Evaluations on four datasets, including our self-collected Small Objects Dataset (SOD), show that MvT seamlessly integrates image-based methods and outperforms existing non-learning and learning trackers in generalized scenarios, and effectively resolves the single-view tracking ambiguities. Being training-free, MvT is applicable when ground-truth annotation is infeasible, thereby highlighting its practical, data-efficient potential. Code is available at https://github.com/zhazhabiu/MvTracking.
Muxi Zha, Banglei Guan, Minzu Liang et al.· IEEE Transactions on Image P...· 0 citations
As a vital component of structural health monitoring, the detection of cracks and water leakage in tunnel linings is essential for ensuring structural durability and operational safety. However, due to complex site conditions, such as non-uniform illumination, surface texture interference, and the slender, blurred nature of defects, traditional manual inspections and threshold-based algorithms often fail to provide reliable damage identification. To address these challenges, this study proposes an end-to-end semantic segmentation framework based on TransUNet. By integrating the local feature extraction of convolutional neural networks (CNNs) with the global dependency modeling of Transformers, the framework significantly enhances the characterization of multi-scale defects and boundary features. A comprehensive dataset comprising public benchmarks and real-world engineering images was developed using a standardized preprocessing and validation pipeline. The proposed method was systematically evaluated against state-of-the-art models like U-Net and DeepLabv3 + . Experimental results demonstrate that the TransUNet framework achieves an IoU of 71.57% for crack segmentation and a Precision of 91.51% for water leakage. Crucially for engineering applications, the geometric error for length and area measurements is maintained within 5%, while the inference latency remains under 200 ms. In terms of precision, boundary preservation, and geometric consistency, the proposed method shows clear advantages over the comparison models, while U‑Net exhibits stronger region overlap for water leakage detection. Overall, the method meets the requirements of offline inspection and near-real-time applications. This data-driven approach provides a robust technical foundation for tunnel defect detection and subsequent maintenance decision-making.
Xinjian Li, Qiaofeng Liu, Gang Yan et al.· PLoS ONE· 0 citations
DECO is proposed, a DEpth-guided CO-visibility reasoning framework for low-altitude UAV visual localization that retains keypoints that are both visually distinctive and geometrically co-visible, improving feature matching and PnP-based pose estimation.
Yi-Bin Ye, Xichao Teng, Shuo Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.