Aug 2026· Intelligent Data Analysis· 0 citations· 2 references
TL;DR
With scratch-trained detector weights and frozen vision-language encoders used only for prompt-side priors, YOLO-TVP achieves competitive prompt-conditioned detection while preserving real-time efficiency, verifying its advances in semantic alignment and prompt-adaptive detection.
Abstract
Open-vocabulary object detection (OVD) models generally leverage vision-language pre-trained models to recognize novel categories via arbitrary text prompts. Nevertheless, their performance is restricted by two core limitations: semantic discontinuity arising from hard binary supervision in contrastive learning, and the inflexibility of single prompts to convey complex detection intents in practical scenarios. To tackle these issues, this paper proposes YOLO-TVP (Text–Visual Prompt), an efficient OVD framework built on the YOLO architecture, with two key designs. First, a Semantic-Target Soft Cross-Entropy (ST-SoftCE) loss is introduced. It constructs semantic target distributions from inter-class similarities in the shared text embedding space for open-vocabulary inference and supervises the detector's classification branch. This design embeds semantic relevance into supervision and enhances fine-grained discrimination among semantically similar categories. Second, a unified class-prompt embedding interface is developed to support both textual and visual prompts. Text prompts are projected into the shared prompt space, while visual prompts are formed by learnable weighted fusion of CLIP semantic priors and backbone visual features, eliminating the need for multi-modal prompt concatenation during inference. Experiments validate the effectiveness: on Flickr30k image-text retrieval with ResNet-50, ST-SoftCE improves Recall@1 by 9.56% over standard cross-entropy. For open-vocabulary detection trained on COCO+Flickr30k and evaluated on LVIS, ST-SoftCE delivers a 13.15% relative mAP
50
improvement to YOLO-World. With scratch-trained detector weights and frozen vision-language encoders used only for prompt-side
priors, YOLO-TVP achieves competitive prompt-conditioned detection while preserving real-time efficiency, verifying its advances in semantic alignment and prompt-adaptive detection.
Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual--semantic cues learned from seen data, in...
Yu-Peng Zhang, Zi-Yi Zhao, Jun-Tao Cheng et al.· 0 citations
Open-world object detection is a key visual support technology for smart-city transportation. Most existing methods are trained on closed-set datasets and recognize only predefined traffic categories. For unknown traffic participants and sudden obstacles in complex scenes, they suffer severe false/missed detection and...
Recent unified open-vocabulary detection (OVD) supports heterogeneous prompts, including text queries, visual exemplars, and their combinations, but often rely on increasingly complex designs such as heavy cross-modal fusion, staged training, and iterative annotation pipelines. We revisit whether such complexity is nec...
Xiao-Yan Wei, Zhi-Min Yao, Rui-Lin Yang et al.· 0 citations
Experimental results show that the proposed open-vocabulary object detection framework performs excellently on multiple benchmark datasets such as LVIS and COCO-O, demonstrating stronger adaptability to complex scenarios.
Tao Liu, Chongwen Wang· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.