Skip to content

PAGE: Towards Practical Human-level Gaze Target Estimation

Jul 2026 · arXiv.org · Vol abs/2607.04860 · 0 citations · 35 references
Computer Science

TL;DR

The architectural improvements and novel training recipe allow PaGE to achieve state-of-the-art performance on several gaze estimation tasks, outperforming humans in 7 out of 9 metrics while reducing the human-AI gap by at least 60% in the remaining 2.

Abstract

Gaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a challenging task that combines high-level understanding of global scene semantics and precise spatial reasoning using human appearance (e.g. pose, eye orientation). As a result, human-level performance remains elusive for existing models, limiting their practical application. To this end, we propose PaGE (Practical Gaze Estimator), a gaze estimation model that explicitly models the complex interaction between scene and head features. Using a PaGE model with a large ViT-H+ backbone as the teacher, we further distill student models with lighter backbones on a much larger and more diverse unlabeled dataset. The architectural improvements and novel training recipe allow PaGE to achieve state-of-the-art performance on several gaze estimation tasks, outperforming humans in 7 out of 9 metrics while reducing the human-AI gap by at least 60% in the remaining 2. The distilled student models retain most of the teacher's performance while being lightweight enough for practical deployment on robots and consumer devices. The code and model checkpoints are available at our project page.

View source

Similar papers

Preprint Sep 2026

From Gaze to Meaning: A Training-Free AI Agent for Unified Grounding and Explanation

This work introduces the first training-free Gaze Target Agent (GTA) for gaze-guided reasoning across tasks such as gaze target prediction, attention localization, and object identification by leveraging pretrained vision-language models, augmenting them with visually guided prompts, and employing a memory-based retrie...

Shayan Nasiriboukani, Sara Atito, Mohammad Nezamipour et al. · 0 citations
Open access Jul 2026

GazeHRNet: Head-Centric Spatial Encoding and Gaze-Aware Feature Interaction for Gaze Target Detection

GazeHRNet is proposed, a head-centric reasoning framework for RGB-based gaze target detection that combines coarse spatial reasoning with fine-grained anisotropic heatmap prediction, enabling reliable target localization under cluttered scenes and varying head positions.

Tianxiang Nan, Chenglizhao Chen, Xi Chen et al. · 0 citations
#small language model Preprint Aug 2026

G3Ego: Gaze-Guided Graphs for Egocentric Action Understanding

G3Ego, a graph-based framework for egocentric action understanding that uses gaze as a structural cue to identify action-relevant entities in the scene, achieves competitive performance compared with video-based approaches and consistently improves Macro-F1 under class-imbalanced evaluation, while avoiding reliance on...

Marko Haralović, Akash Ramakrishnan, E. T. Martínez · 0 citations
Jul 2026

Human pose estimation in library and information science: An exploratory case study of joint attention during virtual storytimes

Human Pose Estimation (HPE), a computer vision technique that detects and tracks human body movements through skeletal keypoints, has been widely used in fields such as healthcare, sports, and education but remains largely unexplored in library and information science (LIS). This exploratory case study examines the use...

Luke LeFebvre, Namjoo Choi, Jerzy W. Jaromczyk et al. · 0 citations
Preprint Sep 2026

GazeDiT: Gaze-Accurate Diffusion Image Generation for Eye Tracking via Spatial Conditioning

Diffusion models are increasingly used to generate synthetic training data, but precise label control remains difficult when the conditioning signal is low-dimensional and coarse. Text-conditioned images are judged by broad prompt consistency, whereas supervised training requires precise correspondence between each ima...

Dong-Ze Wu, David Colmenares, Feng-Ting Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.