Skip to content

V-Align: Visual Forced Alignment via Phoneme to Video Optimal Path Traversal

· 0 citations · 16 references

TL;DR

V-Align is proposed, a structured framework that formulates VFA as an optimal mono-tonic path traversal problem over a frame–phoneme compatibility lattice and learns a soft traversal posterior and recovers discrete boundaries via dynamic programming.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

VETO: Video Efficient Token Optimization for Vision Language Models

Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We...

Gueter Josmy Faure, Hao-Ping Wang, Min-Hung Chen et al. · 0 citations
Preprint Sep 2026

DSE-VTG: Dual-Side Enhancement for Training-Free Video Temporal Grounding

Text-guided Video Temporal Grounding (VTG) aims to localize the relevant segments in an untrimmed video based on text queries, yet collecting dense temporal annotations and training task-specific models remain costly and brittle under distribution shift. Recent training-free VTG approaches mitigate this issue by direct...

Zhuo Cao, Bing-Qing Zhang, Sen Wang et al. · 0 citations
Preprint Sep 2026

TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models

Visual grounding maps language referents to spatial targets and is central to open-vocabulary perception with vision-language models. Existing methods have made substantial progress on single-frame and video-based visual grounding, yet under streaming inputs they still suffer from identity drift, cross-frame inconsiste...

Le-Qian Ding, Jun-Ning Qiu, Man-Wen Yang et al. · 0 citations
Preprint Sep 2026

Semantic Refinement of Universal Audio Representations through Audio-Description Alignment

Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of...

Le-Jun Min, Jun-Yu Dai, Rui-Chen Zheng et al. · 0 citations
Open access 2026

Cross-Modal Temporal Alignment for Action Grounding in Videos

A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.

Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.