Skip to content

Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning

Jul 2026 · arXiv.org · Vol abs/2607.23077 · 0 citations · 52 references
Computer Science

TL;DR

This work decomposes multi-entity temporal dynamics into three interaction types: spatial interactions among entities, temporal interactions across time, and cross interactions coupling the two domains, and proposes a structured spatio-temporal transformer block that explicitly models all three within a single stage.

Abstract

Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions. Transformer-based approaches perform well but often rely on deep stacks of layers to learn these heterogeneous dependencies implicitly, increasing computational cost. We revisit this problem from a structural perspective and decompose multi-entity temporal dynamics into three interaction types: spatial interactions among entities, temporal interactions across time, and cross interactions coupling the two domains. We propose a structured spatio-temporal transformer block that explicitly models all three within a single stage. It uses parallel spatial and temporal self-attention, followed by bidirectional cross-attention, and combines the outputs through learnable gated fusion. By directly encoding these complementary views, the model reduces the need for deep stacking. We evaluate the approach on video-based group activity recognition, skeleton-based human interaction analysis, and wearable sensor-based activity recognition. Despite its simplicity, the single structured Transformer block matches or outperforms deeper architectures with only 1.76M parameters. The results suggest that depth in prior models partly compensates for implicit and entangled interaction modeling, whereas explicit factorization offers a more efficient and transparent alternative. More broadly, this work supports a structure-first design principle: expressive multi-entity temporal reasoning can emerge by exposing interaction structure rather than relying on depth.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Event Interaction in Low-Rank Bottlenecks for Temporal Relation Extraction

Temporal relation extraction determines whether an event occurs before, after, or simultaneously with another event, and therefore relies on accurately modeling how the two events interact. Mainstream systems achieve this by concatenating event spans or using shallow fusion, which works well when all model parameters a...

Wei Sun, Tingyu Qu, Jesse Davis et al. · 0 citations
#machine learning Preprint Sep 2026

Temporal Heterogeneous Graph Pretraining for Relational Deep Learning

Relational deep learning models database rows and foreign-key links as a heterogeneous graph for prediction from record attributes and relational context. These graphs contain two distinct temporal signals: record age changes with the prediction cutoff, while intervals between observed records remain fixed. Prior work...

Yi-Xin Peng, Er Jin, Diego Collarana et al. · 0 citations
Preprint Aug 2026

Topology-Masked Unified Backbone for Joint Feature Interaction and Multi-Domain Sequence Modeling

Large-scale post-click conversion rate (CVR) prediction requires jointly modeling heterogeneous feature interactions and dependencies over multi-domain user behavior sequences. Existing industrial ranking models usually handle these two aspects with separate modules. Recent unified architectures attempt to incorporate...

Zhiwu Zhu, Dezheng Han, Jingjie Xia et al. · 0 citations
Open access 2026

Encoder-Level Temporal Fusion Transformers for Robust Multi-Object Tracking

Experiments on three challenging benchmarks show that GTF improves identity association and overall tracking robustness, particularly under fast motion, appearance ambiguity, and frequent occlusions, and ablation studies further support the roles of encoder-level temporal fusion, adaptive gating, entropy regularization...

Jinho Kim, Kuk-Jin Yoon · 0 citations
Sep 2026

Spectral-temporal dual view graph convolutional networks: A deep learning framework for dynamic bipartite graphs.

Dynamic bipartite graphs (DBGs) are widely used in real-world scenarios, where representation learning is particularly challenging due to the heterogeneity of node types and the temporal evolution of interactions. A key difficulty lies in jointly capturing non-stationary micro-level preference dynamics and macro-level...

Zhe-Zhe Xing, Yu-Xin Ye, Zi-Heng Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.