The results establish explicit-implicit matching as an effective and efficient mechanism for spatio-temporal instance interaction and Integrating temporal MatchFusion into SparseDrive further improves perception within an E2E framework without additional supervision.
Abstract
Sparse instance representations provide a compact interface for spatial LiDAR-camera and temporal past-current interaction in multimodal perception and E2EAD. Effective interaction requires reliable instance correspondences despite geometric discrepancies and heterogeneous semantic representations. Attention-based methods exploit contextual semantics but often require specialized representation alignment, increasing computational overhead. In contrast, association based on structured object states is efficient and interpretable but lacks contextual evidence to resolve ambiguous matches. To combine these complementary strengths, we propose MatchFusion, a learnable instance matching and fusion module for spatio-temporal multimodal autonomous driving. MatchFusion initializes pairwise affinities using geometric similarity and category consistency, then selectively refines structurally plausible associations using instance embeddings. The resulting soft matchmap guides a common residual aggregation operator for adaptive information exchange. This unified matching-fusion formulation supports spatial LiDAR-camera and temporal past-current interaction, using multi-view image-plane geometry and motion-compensated BEV geometry as the respective structural priors. Experiments on nuScenes demonstrate consistent perception gains across diverse front-end configurations. Compared with a prior instance-centric fusion method, the MatchFusion-equipped system achieves higher perception accuracy while reducing FLOPs by 55.3% and GPU memory usage by 39.3%, with the matching-fusion module accounting for only 3.7% of total perception latency. Integrating temporal MatchFusion into SparseDrive further improves perception within an E2E framework without additional supervision. These results establish explicit-implicit matching as an effective and efficient mechanism for spatio-temporal instance interaction.
Multimodal learning aims to integrate heterogeneous observations such as images, text, depth, and radar to improve perception and reasoning. However, most existing multimodal models implicitly assume that cross-modal observations are well aligned, an assumption that rarely holds in real-world scenarios due to viewpoint...
Qian Li, Ya-Heng Wang, Qian Huang et al.· IEEE Transactions on Pattern...· 0 citations
Achieving accurate LiDAR-based place recognition is a crucial step towards reliable autonomous navigation, as it enables robust loop closure and global localization without GPS. However, existing 3D point cloud descriptors often give up fine local geometry for broad semantic context. We propose SBC-Net, a dual-branch a...
Minseong Park, DoHyeong Kwon, Jae-Jin Jeon et al.· IEEE Access· 0 citations
Long-horizon robotic manipulation requires a policy to bridge task-level semantic reasoning with metric three-dimensional interaction geometry. Existing vision–language–action policies usually acquire geometry implicitly from visual tokens or introduce deterministic intermediate variables only in the image plane, which...
Li Lin, Ming-Hao Shi, Teng-Long Wang· Applied Informatics· 0 citations
Modeling spatiotemporal coupling is a key challenge in building physical intelligence across scales, from microscopic to macroscopic. Existing models capture such structure broadly through physics-motivated dynamical formulations or learning-motivated architectures. The former provide stronger priors but may constrain...
Yu-Hao Li, Louie Hong Yao, Tian-Yi Shi et al.· 0 citations
Comparing current and past observations helps robots understand scene changes and select subsequent actions during manipulation. However, comparing visual features at the same image location can mix different scene content when objects or the camera move. We introduce CoRe-WAM, a world-action model that incorporates co...
Bing-Heng Zhou, Jia-Long Liu, Jia-Nan Wang et al.· 0 citations
World4V2X is proposed, the first world model framework tailored for V2X cooperative perception, which introduces a spatial observability modeling module that defines spatial consistency boundaries to distinguish reliable regions from uncertain ones, thereby enabling spatial consistency modeling over multi-agent heterog...
Rui Wang, Shuai Wang, Xiang-Yi Qin et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.