Skip to content

Interpretable Temporal Video Reasoning with EventGraph and EventField

Sep 2026 · 0 citations · 16 references
Computer Science

TL;DR

The results indicate that structured temporal representations can support both performance and inspectability by preserving symbolic structure, capturing temporal continuity, and providing human-readable diagnostics for video reasoning.

Abstract

We present a structured temporal video reasoning pipeline built around a discrete EventGraph, a continuous EventField, and a human-readable EventGlyph view. On a calibrated EPIC-KITCHENS subset of 10 videos and 50 temporal reasoning questions, EventField+Glyph achieves 0.98 overall accuracy, which is higher than the caption baseline by +0.40 (paired p = 1.1 \times 10^{-5}) and direct VLM-only QA by +0.20 (p = 0.0063) on this subset. We further evaluate annotation-source variations, including manual, heuristic, and heuristic+Gemini pipelines, and find that the best structured method stays above the caption baseline across settings. We also include cross-video pair benchmarking and an appendix gallery of glyph outputs for all studied videos. Overall, the results indicate that structured temporal representations can support both performance and inspectability by preserving symbolic structure, capturing temporal continuity, and providing human-readable diagnostics for video reasoning.

View source

Similar papers

Preprint Aug 2026

Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting

Structured video prompting is introduced, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding.

Sadegh Mohammadian · 0 citations
Aug 2026

DiaVTG: multi-turn reasoning framework for video temporal grounding

DiaVTG, a novel VTG framework designed to enhance temporal localization precision, is proposed to reformulate temporal localization as a video understanding problem and demonstrates that the training-free method consistently improves performance across various Vid-LLM architectures.

Hong-Yu Huang, Junyi Yang, Sipeng Yang et al. · 0 citations
Review Sep 2026

Event-Grounded Football News Generation from Match Videos with Parameter-Efficient Large Language Models

Automated football news generation from raw videos requires bridging spatiotemporal perception with factual text composition. This study develops an end-to-end, event-based framework that converts match videos into fact-grounded reports. The framework uses an Inflated Three-Dimensional ConvNet (I3D) backbone with multi...

Yi-Feng Wang, Yi-Hang Huang · 0 citations
#artificial intelligence Preprint Sep 2026

AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

Findings show that semantic reference-text overlap should not be treated as a proxy for temporal localization, and that AVTrace can identify task-specific weaknesses while providing a testbed for temporal post-training.

Long-Yin Zhang, P. Mahendra, Chengwei Wei et al. · 1 citation
Preprint Sep 2026

The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs

VT-Contrast supervises selected late-layer last-frame video tokens, where temporal information is expected to be integrated before language generation, and contrasts order-preserving views with same-video reordered counterfactuals graded by Kendall tau distance.

Yu-Meng Shi, Quan-Yu Long, Yin Wu et al. · 0 citations
Preprint Aug 2026

VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

VideoVIBE is introduced, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks and V2Lens is proposed, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-...

Jia-Jun Xu, Yang-Hao Zhou, Jing Liao et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.