Skip to content
Conference Open access

Bringing Real-World Relations into Video Generation with Graph-Structured Knowledge

2026 · Annual Meeting of the Association for Computational Linguistics · pp. 3756-3771 · 0 citations · 35 references
Computer Science

TL;DR

Comprehensive evaluations demonstrate that incorporating graph-structured knowledge significantly enhances compositionality and the accurate portrayal of real-world interactions in generated videos.

Abstract

Recent proprietary video generation models have demonstrated remarkable proficiency in synthesizing highly realistic videos from textual instructions. Most open-source text-to-video models, however, still struggle to accurately simulate real-world physics and dynamic entity interactions. Existing approaches rely on scaling laws and large-scale, high-quality video datasets to implicitly learn physical dynamics, yet this paradigm is constrained by prohibitive costs and the burdensome demands of data curation. Motivated by this, we propose a novel framework that integrates graph-structured temporal knowledge into video latent diffusion models to enhance compositional generation and interaction fidelity. Our framework constructs video scene graphs specifically designed to capture entity relationships, temporal dynamics, and global scene context. These graph-structured representations guide the generation process through cross-attention mechanisms. Additionally, we introduce Graph-Aligned De-noising Loss (GADL), a training objective that ensures adherence to conditioned graphs by incorporating node modification tasks within the denoising process, leveraging synchronized edited video-graph pairs. Comprehensive evaluations demonstrate that incorporating graph-structured knowledge significantly enhances compositionality and the accurate portrayal of real-world interactions in generated videos.

Read PDF

Similar papers

Open access Aug 2026

Long-Horizon Video Generation with Temporally Consistent Diffusion and Scene-Graph Guidance

A novel framework that integrates temporally consistent diffusion models with dynamic scene-graph guidance that structurally constrains the generative process, ensuring that objects, their attributes, and their interrelationships remain stable over extended durations is introduced.

Jacob A. Jenkins · 0 citations
Jul 2026

GraphVid: Interactive Graph-Controllable Video Generation

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scal...

Vedant r Shah, O. Susladkar, Tushar Prakash et al. · 0 citations
#large language models Preprint Sep 2026

ViTAL-X: Video-Text Alignment with Cross-Modal Temporal Edits

ViTAL-X, a lightweight model that equips frozen image-text backbones with temporal awareness while preserving their foundational spatial knowledge, achieves state-of-the-art performance and demonstrates that targeted, high-quality temporal alignment provides a highly efficient alternative to pure scaling.

V. SethuramanT., Savya Khosla, O. Susladkar et al. · 0 citations
Preprint Aug 2026

Modeling Scientific Experiment Scenes: Dataset and Model

The Cross-Modal Dual-Path Generator (CM-DPG), a model for robust open-vocabulary SGG that enhances object-level semantic representations through joint visual-textual encoding and improves relational reasoning using complementary visual and geometric cues, is proposed.

Ming-Hao Zou, Qingtian Zeng, Shangkun Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models...

Ruibo Ming, Lei Sun, De-Heng Zhang et al. · 0 citations

Predictive Structure Improves Video Diffusion Dynamics

Experiments show that LDO substantially improves physical commonsense, object permanence, and trajectory fidelity while preserving visual quality, suggesting that predictive latent supervision offers a practical route to make video generators not only photorealistic but also physically legible.

Lu-Jun Mi, Yujia Chen, Alex Dimakis et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.