V-Align is proposed, a structured framework that formulates VFA as an optimal mono-tonic path traversal problem over a frame–phoneme compatibility lattice and learns a soft traversal posterior and recovers discrete boundaries via dynamic programming.
Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We...
Text-guided Video Temporal Grounding (VTG) aims to localize the relevant segments in an untrimmed video based on text queries, yet collecting dense temporal annotations and training task-specific models remain costly and brittle under distribution shift. Recent training-free VTG approaches mitigate this issue by direct...
Zhuo Cao, Bing-Qing Zhang, Sen Wang et al.· 0 citations
Visual grounding maps language referents to spatial targets and is central to open-vocabulary perception with vision-language models. Existing methods have made substantial progress on single-frame and video-based visual grounding, yet under streaming inputs they still suffer from identity drift, cross-frame inconsiste...
Le-Qian Ding, Jun-Ning Qiu, Man-Wen Yang et al.· 0 citations
Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of...
Le-Jun Min, Jun-Yu Dai, Rui-Chen Zheng et al.· 0 citations
Inspired by the linguistic phenomenon of code-switching, MMCS interleaves vision and language by replacing textual entities with their corresponding visual objects, enforcing local vision-language grounding.
Changhao Xiang, Shangyu Xing, Zhen Wu et al.· 0 citations
A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.
Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.