Skip to content

Author

Minghao Liu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

SRCDet: sparse fusion of surround-view radar and camera for 3D object detection

Multimodal fusion of cameras and millimeter-wave radars is critical for robust all-weather object detection and ensuring vehicle safety in real-world autonomous driving vehicles. However, existing radar-camera fusion methods that rely on a unified Bird’s Eye View (BEV) representation often suffer from information loss and limited cross-modal interaction. To address these limitations, a query-based multimodal fusion framework, termed SRCDet, is proposed for camera-4D radar fusion. The framework processes features in parallel across both BEV and Perspective View spaces, where deformable attention is employed to achieve dynamic cross-view alignment. By integrating radar attribute features, a local–global dual-branch query generation mechanism is designed to produce high-quality 3D detection proposals. Furthermore, a graph neural network-based cross-fusion module is introduced to model complex inter-feature relationships through a heterogeneous interaction graph. Extensive experiments on the OmniHD-Scenes and NuScenes datasets demonstrate that SRCDet achieves consistent improvements across nearly all metrics and low error rates in clear and adverse weather conditions, highlighting its practical adaptability to automotive-grade systems and effectiveness in safety-critical real-world autonomous driving scenarios.

Wen-Jin Ai, Lianqing Zheng, Long Yang et al. · 0 citations
Preprint Aug 2026

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.

Sirun Li, Minghao Liu, Ling Dai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.