Skip to content

Author

Pengfei Zhou

We have 4 of 8 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

EgoTSR++: Egocentric Spatiotemporal Reasoning for Task Progress Understanding

Vision-Language Models (VLMs) have advanced rapidly in static visual understanding, yet remain unreliable when judging how an egocentric task is progressing. Given a task instruction and two visual observations, a model should determine which state is closer to the goal by analyzing task-relevant object configurations...

Xiao-Da Yang, Can Wang, Yu-Xiang Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models

Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded bench...

Hao Wang, Tao Yu, Liu-Zhou Zhang et al. · 0 citations
Preprint Sep 2026

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whos...

AgiBot Research Team, Renhang Liu, Wen-Zhi Zhao et al. · 2 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.