Skip to content
Preprint

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

Aug 2026 · 0 citations · 56 references
Computer Science

TL;DR

360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning, and evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level.

Abstract

We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.

View source

Similar papers

NavAI: an application-agnostic llm framework for navigation tasks in virtual reality environments

NavAI is an extensible navigation framework that leverages large language models (LLMs) to support both basic action commands and multi-step goal-oriented navigation through an application-agnostic screenshot-and-control interface and explores optimization strategies for virtual scene understanding and navigation goal...

Jiajie Wang, Sumesh Surendran Letha, M. DiGiovanni et al. · 0 citations
Preprint Aug 2026

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

This paper proposes UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data, and hopes it will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.

Tianjie Ju, Zheng Wu, Yueqing Sun et al. · 1 citation
Open access Aug 2026

Embodied Scene Rearrangement Planning

Experimental results demonstrate that current methods struggle to complete the ESRP task efficiently, highlighting ESRP as a challenging frontier for embodied agents in scene understanding and long-horizon task planning.

Can-Zhi Chen, Zan Wang, Siqi Zhu et al. · 0 citations
Open access Sep 2026

A visual perception-guided framework for camera path planning in large-scale digital twin and AI-generated 3D scenes

This study investigates visual, cognitive, and spatial saliency indicators of architectural landmarks in large-scale ancient city ruins and digital twin virtual environments. It further examines users’ cognitive demands during pathfinding in large, complex environments and explores how different exploration strategies...

Yan Zhang, Wei Chen, Gang Yang · 0 citations
Oct 2026

ForexNav: Foresight Exploratory Navigation in Complex and Unknown Indoor Environments

Autonomous navigation in unknown, complex indoor environments remains challenging due to limited sensing range and severe partial observability. Conventional methods rely on local maps without foresight, causing dead-ends and long detours, while local goal selection based on Euclidean distance or frontier coverage fail...

Hong-Yu Song, Yun-Fang Ren, Ji-Gui Miao et al. · 0 citations
#computer vision Preprint Aug 2026

GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation

This work introduces GeoAgent, an agentic environment-based benchmark that requires agents to navigate Street View environments to refine their geolocalization through sequential reasoning, and establishes the challenges of embodied navigation and geospatial reasoning.

Arka Mukherjee, Soham Roy, Kartikeya Trivedi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.