Vorch-Omni is presented, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation that supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video tr...
Vorch Team, Xiaoyu Chen, Yang Ding et al.· 0 citations
Agentic Real2Sim is introduced, a framework for generalized physical world modeling with vision-language agents, converting a real-world recording of object-robot interaction into a simulatable episodic twin which preserves observations, geometries, robot interactions, and object states.
This work introduces the cross-predictive JEPA (JEPA-x), which grounds latent dynamics in privileged physical trajectories, and shows that direct physical-state regression improves decodability without improving forecastability or control, indicating that the benefit comes from shaping latent dynamics rather than merel...
Kehan Wen, Ziming Li, Si-Yuan Luo et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.