Skip to content

YOLO-TVP: Real-time open-vocabulary object detection with Semantic-Target Soft Cross-Entropy and text-visual prompts

Aug 2026 · Intelligent Data Analysis · 0 citations · 2 references

TL;DR

With scratch-trained detector weights and frozen vision-language encoders used only for prompt-side priors, YOLO-TVP achieves competitive prompt-conditioned detection while preserving real-time efficiency, verifying its advances in semantic alignment and prompt-adaptive detection.

Abstract

Open-vocabulary object detection (OVD) models generally leverage vision-language pre-trained models to recognize novel categories via arbitrary text prompts. Nevertheless, their performance is restricted by two core limitations: semantic discontinuity arising from hard binary supervision in contrastive learning, and the inflexibility of single prompts to convey complex detection intents in practical scenarios. To tackle these issues, this paper proposes YOLO-TVP (Text–Visual Prompt), an efficient OVD framework built on the YOLO architecture, with two key designs. First, a Semantic-Target Soft Cross-Entropy (ST-SoftCE) loss is introduced. It constructs semantic target distributions from inter-class similarities in the shared text embedding space for open-vocabulary inference and supervises the detector's classification branch. This design embeds semantic relevance into supervision and enhances fine-grained discrimination among semantically similar categories. Second, a unified class-prompt embedding interface is developed to support both textual and visual prompts. Text prompts are projected into the shared prompt space, while visual prompts are formed by learnable weighted fusion of CLIP semantic priors and backbone visual features, eliminating the need for multi-modal prompt concatenation during inference. Experiments validate the effectiveness: on Flickr30k image-text retrieval with ResNet-50, ST-SoftCE improves Recall@1 by 9.56% over standard cross-entropy. For open-vocabulary detection trained on COCO+Flickr30k and evaluated on LVIS, ST-SoftCE delivers a 13.15% relative mAP 50 improvement to YOLO-World. With scratch-trained detector weights and frozen vision-language encoders used only for prompt-side priors, YOLO-TVP achieves competitive prompt-conditioned detection while preserving real-time efficiency, verifying its advances in semantic alignment and prompt-adaptive detection.

View source

Similar papers

Preprint Oct 2026

RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection

Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual--semantic cues learned from seen data, in...

Yu-Peng Zhang, Zi-Yi Zhao, Jun-Tao Cheng et al. · 0 citations
Open access Oct 2026

SAVPEF-Net: Symmetric Semantic–Visual Prompt Fusion for Fine-Grained Open-World Object Detection in Intelligent Transportation

Open-world object detection is a key visual support technology for smart-city transportation. Most existing methods are trained on closed-set datasets and recognize only predefined traffic categories. For unknown traffic participants and sudden obstacles in complex scenes, they suffer severe false/missed detection and...

Dou-Ping Bai, Zheng-Biao Jing, Dong-Ling Jing · 0 citations
Preprint Aug 2026

OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection

Recent unified open-vocabulary detection (OVD) supports heterogeneous prompts, including text queries, visual exemplars, and their combinations, but often rely on increasingly complex designs such as heavy cross-modal fusion, staged training, and iterative annotation pipelines. We revisit whether such complexity is nec...

Xiao-Yan Wei, Zhi-Min Yao, Rui-Lin Yang et al. · 0 citations
Conference Aug 2026

Open-vocabulary object detection based on fine-grained alignment

Experimental results show that the proposed open-vocabulary object detection framework performs excellently on multiple benchmark datasets such as LVIS and COCO-O, demonstrating stronger adaptability to complex scenarios.

Tao Liu, Chongwen Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.