The architectural improvements and novel training recipe allow PaGE to achieve state-of-the-art performance on several gaze estimation tasks, outperforming humans in 7 out of 9 metrics while reducing the human-AI gap by at least 60% in the remaining 2.
Abstract
Gaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a challenging task that combines high-level understanding of global scene semantics and precise spatial reasoning using human appearance (e.g. pose, eye orientation). As a result, human-level performance remains elusive for existing models, limiting their practical application. To this end, we propose PaGE (Practical Gaze Estimator), a gaze estimation model that explicitly models the complex interaction between scene and head features. Using a PaGE model with a large ViT-H+ backbone as the teacher, we further distill student models with lighter backbones on a much larger and more diverse unlabeled dataset. The architectural improvements and novel training recipe allow PaGE to achieve state-of-the-art performance on several gaze estimation tasks, outperforming humans in 7 out of 9 metrics while reducing the human-AI gap by at least 60% in the remaining 2. The distilled student models retain most of the teacher's performance while being lightweight enough for practical deployment on robots and consumer devices. The code and model checkpoints are available at our project page.
This work introduces the first training-free Gaze Target Agent (GTA) for gaze-guided reasoning across tasks such as gaze target prediction, attention localization, and object identification by leveraging pretrained vision-language models, augmenting them with visually guided prompts, and employing a memory-based retrie...
Shayan Nasiriboukani, Sara Atito, Mohammad Nezamipour et al.· 0 citations
GazeHRNet is proposed, a head-centric reasoning framework for RGB-based gaze target detection that combines coarse spatial reasoning with fine-grained anisotropic heatmap prediction, enabling reliable target localization under cluttered scenes and varying head positions.
Tianxiang Nan, Chenglizhao Chen, Xi Chen et al.· Italian National Conference...· 0 citations
G3Ego, a graph-based framework for egocentric action understanding that uses gaze as a structural cue to identify action-relevant entities in the scene, achieves competitive performance compared with video-based approaches and consistently improves Macro-F1 under class-imbalanced evaluation, while avoiding reliance on...
Marko Haralović, Akash Ramakrishnan, E. T. Martínez· 0 citations
Human Pose Estimation (HPE), a computer vision technique that detects and tracks human body movements through skeletal keypoints, has been widely used in fields such as healthcare, sports, and education but remains largely unexplored in library and information science (LIS). This exploratory case study examines the use...
Luke LeFebvre, Namjoo Choi, Jerzy W. Jaromczyk et al.· Journal of Librarianship and...· 0 citations
Diffusion models are increasingly used to generate synthetic training data, but precise label control remains difficult when the conditioning signal is low-dimensional and coarse. Text-conditioned images are judged by broad prompt consistency, whereas supervised training requires precise correspondence between each ima...
Dong-Ze Wu, David Colmenares, Feng-Ting Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.