The results indicate that structured temporal representations can support both performance and inspectability by preserving symbolic structure, capturing temporal continuity, and providing human-readable diagnostics for video reasoning.
Abstract
We present a structured temporal video reasoning pipeline built around a discrete EventGraph, a continuous EventField, and a human-readable EventGlyph view. On a calibrated EPIC-KITCHENS subset of 10 videos and 50 temporal reasoning questions, EventField+Glyph achieves 0.98 overall accuracy, which is higher than the caption baseline by +0.40 (paired p = 1.1 \times 10^{-5}) and direct VLM-only QA by +0.20 (p = 0.0063) on this subset. We further evaluate annotation-source variations, including manual, heuristic, and heuristic+Gemini pipelines, and find that the best structured method stays above the caption baseline across settings. We also include cross-video pair benchmarking and an appendix gallery of glyph outputs for all studied videos. Overall, the results indicate that structured temporal representations can support both performance and inspectability by preserving symbolic structure, capturing temporal continuity, and providing human-readable diagnostics for video reasoning.
Structured video prompting is introduced, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding.
DiaVTG, a novel VTG framework designed to enhance temporal localization precision, is proposed to reformulate temporal localization as a video understanding problem and demonstrates that the training-free method consistently improves performance across various Vid-LLM architectures.
Hong-Yu Huang, Junyi Yang, Sipeng Yang et al.· The Visual Computer· 0 citations
Automated football news generation from raw videos requires bridging spatiotemporal perception with factual text composition. This study develops an end-to-end, event-based framework that converts match videos into fact-grounded reports. The framework uses an Inflated Three-Dimensional ConvNet (I3D) backbone with multi...
Yi-Feng Wang, Yi-Hang Huang· International journal of pat...· 0 citations
Findings show that semantic reference-text overlap should not be treated as a proxy for temporal localization, and that AVTrace can identify task-specific weaknesses while providing a testbed for temporal post-training.
Long-Yin Zhang, P. Mahendra, Chengwei Wei et al.· 1 citation
VT-Contrast supervises selected late-layer last-frame video tokens, where temporal information is expected to be integrated before language generation, and contrasts order-preserving views with same-video reordered counterfactuals graded by Kendall tau distance.
Yu-Meng Shi, Quan-Yu Long, Yin Wu et al.· 0 citations
VideoVIBE is introduced, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks and V2Lens is proposed, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-...
Jia-Jun Xu, Yang-Hao Zhou, Jing Liao et al.· 1 citation
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.