Skip to content
Open access

BMF-DETR: Pseudo-Depth-Guided Bidirectional Multi-Strategy Fusion for End-to-End Object Detection

Sep 2026 · Applied Informatics · 0 citations · 10 references

Abstract

Transformer-based detectors model long-range context effectively, yet their representations remain dominated by RGB appearance and can become unreliable in cluttered, occluded, or crowded scenes. We present BMF-DETR, a pseudo-depth-guided detector that introduces RGB-derived geometric structure without requiring a depth sensor. DA3Mono-Large from Depth Anything 3 generates spatially aligned pseudo-depth maps offline, while two ResNet-50 streams encode appearance and relative geometry. Bidirectional cross-modal attention (BCMA) establishes two-way correspondence, and multi-strategy fusion (MSF) combines the streams through global calibration, channel allocation, and spatial gating before squeeze-and-excitation (SE) recalibration. On the fixed validation/evaluation split of the 2024 Roboflow-curated PASCAL VOC derivative, the complete model reaches 60.80 AP, compared with 52.80 AP for a capacity-matched dual-RGB control. BMF-DETR obtains 49.30 AP on COCO 2017. Its detector contains 58 M parameters and requires 103 GFLOPs; these figures exclude offline pseudo-depth generation. A shared-low-level variant retains 60.10 AP with 50 M parameters and 87 GFLOPs. The results show that pseudo-depth can serve as a useful auxiliary representation when its contribution is separated from capacity effects and evaluated under controlled fusion settings.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.