Skip to content
Preprint

Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

Motion-based tokenization provides a compact representation for event-aligned egocentric gaze streams, while the evaluation identifies how target predictability and event construction shape cross-dataset conclusions.

Abstract

Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw traces preserve detail but are noisy and device-dependent, while coarse event labels are easy to model but can discard local motion structure. We formulate event-aligned, fixed-horizon angular displacement as an interpretable, event-conditioned motion vocabulary and compare it with event-only, spatial, absolute-angle, learned vector-quantized, and continuous representations. To assess transfer alongside target predictability and token collapse, our evaluation combines next-token prediction with target-domain regret, low-order target references, paired bootstrap, order sensitivity, motif overlap, and frozen structural probes. In an event-aligned headset benchmark, angular-motion tokens have lower target-domain regret than frozen-codebook VQ tokens in one transfer direction, while the reverse direction is inconclusive. The probes reveal complementary representation properties, and event-only tokens show that low perplexity can retain little motion information. On a third egocentric dataset, a matched comparison of I-VT, native, and frame-span interfaces shows that event construction materially changes transfer: native events have the lowest regret into EGTEA, while frame-span events have zero motif overlap and fail severely as a source. Motion-based tokenization therefore provides a compact representation for event-aligned egocentric gaze streams, while the evaluation identifies how target predictability and event construction shape cross-dataset conclusions.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Diffusion models for eye-gaze trajectory generation using position and velocity representations

Eye-tracking data are expensive to collect, requiring specialized hardware and controlled laboratory conditions, and difficult to share because of privacy constraints. We address this using two complementary denoising diffusion probabilistic models (DDPMs) for unconditional generation of eye-gaze dynamics from visual-s...

Laxman Basnet, Alexander Szorkovszky, Pedro G. Lind et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

Time-Frequency Geometric Cross-Attention (TFGCA), a drop-in module repairing both blind spots in vision-language-action policies, uses a per-dimension learnable stationary wavelet transform to decompose the action chunk into time-frequency tokens, and each time token retrieves information from them via a cross-attentio...

Sheng-Ye Dong, Hao-Chen Niu, Hao Liu et al. · 0 citations
Preprint Aug 2026

VidParse: Online Parsing of Egocentric Procedures Like a Pro

VidParse is presented, an online, training-free framework that treats activity understanding as a graph-constrained inference problem and achieves up to a 10x improvement in complex multi-step parsing accuracy over strong online baselines, all without requiring a single gradient update.

Anubhav Gupta, A. Kambhamettu, Vatsal Agarwal et al. · 0 citations
Preprint Aug 2026

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged inputs such as video, depth, or CAD models, or rec...

T. Ding, Zhen Luo, Yi-Xuan Yang et al. · 0 citations
#small language model Preprint Aug 2026

G3Ego: Gaze-Guided Graphs for Egocentric Action Understanding

G3Ego, a graph-based framework for egocentric action understanding that uses gaze as a structural cue to identify action-relevant entities in the scene, achieves competitive performance compared with video-based approaches and consistently improves Macro-F1 under class-imbalanced evaluation, while avoiding reliance on...

Marko Haralović, Akash Ramakrishnan, E. T. Martínez · 0 citations
Preprint Sep 2026

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization, is introduced, suggesting its potential to mitigate performance degradation when training on large and diverse dataset mixtures.

Yu-Fei Duan, Hang Yin, A. Longhini et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.