WildCity, a real-world multimodal dataset collected by autonomous fleets traversing complex urban environments, aims to catalyze progress not only in city-scale rendering, but more broadly in the pursuit of AI that can perceive, remember, and reason across space at a scale comparable to human cognition.
Abstract
Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilometers. Can AI build spatial representations at a comparable scale? Although recent foundation models have advanced scene reconstruction and embodied intelligence, scaling to entire cities remains an open challenge, primarily due to the lack of city-scale data. To bridge the gap, we introduce WildCity, a real-world multimodal dataset collected by autonomous fleets traversing complex urban environments. Our dataset includes 18 trajectories, each averaging 83.7 kilometers in length, and preserves the core challenges of in-the-wild perception, e.g., dynamic objects, lighting variations, and imperfect camera poses. We further establish an urban-tailored reconstruction baseline and convert the reconstructed environments into a closed-loop simulator. Beyond the dataset and baseline, we systematically analyze the key challenges on the path to simulation-ready urban digital twins: scalability, extrapolation, and uncertainty. Ultimately, WildCity aims to catalyze progress not only in city-scale rendering, but more broadly in the pursuit of AI that can perceive, remember, and reason across space at a scale comparable to human cognition. Project page: https://han-xiangyu.github.io/Wild-City/
360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning, and evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level.
Kenta Watanabe, Atsuyuki Miyai, Mizuki Takenawa et al.· 1 citation
This paper proposes UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data, and hopes it will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.
Tianjie Ju, Zheng Wu, Yue-Qing Sun et al.· 1 citation
WorldClaw is presented, a fully agentic, coarse-to-fine framework for open-world 3D scene generation that produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.
Chunchao Guo, Jinpeng Li, Yang Li et al.· 5 citations
RoadWeaver is presented, a coarse-to-fine framework for from-scratch generation of diverse, large-scale HD maps, which first synthesizes a global road layout, expands it into a connected road network, and then constructs lane-level geometry with topologically consistent lane connectivity.
Yue-Yuan Li, Zexi Chen, Weijie Xi et al.· 1 citation
HELIOS is proposed, a novel image relighting approach that relies on unlabeled real-world datasets without requiring any paired images for training, and integrates albedo-based conditioning into a cycle-consistent diffusion pipeline to prevent identity collapse and ensure accurate domain translation.
Hala Djeghim, Nathan Piasco, Luis Roldão et al.· 0 citations
This work introduces GeoAgent, an agentic environment-based benchmark that requires agents to navigate Street View environments to refine their geolocalization through sequential reasoning, and establishes the challenges of embodied navigation and geospatial reasoning.
Arka Mukherjee, Soham Roy, Kartikeya Trivedi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.