Generative world-action models (WAMs) jointly generate future video and vehicle actions, while their action branches remain primarily optimized by expert imitation. Yet imitation provides no explicit closed-loop geometric verdict for generated trajectories, making verification important during both training and deploym...
Feng-Cheng Yu, Dhruv Parikh, Jun-Jie Ye et al.· 0 citations
World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric...
Dhruv Parikh, Feng-Cheng Yu, Quan-Kai Gao et al.· 0 citations
FOLIO is introduced, a training-free focused semantic memory system that records important parts of the stream in higher detail while keeping surrounding context compactly and substantially reducing the cost of maintaining streaming memory by reserving detailed records for focused entities and storing surrounding conte...