Reinforcement learning (RL) for dexterous manipulation must discover finger-object contacts and then control the object precisely; the action noise that serves the first goal can interfere with the second. In trajectory-guided settings such as ViViDex, where RL refine hand-object trajectories from human video, our base...
Hao-Yu Wang, Si-Yuan Qian, Yan-Jun Li et al.· 0 citations
Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on an unverified assumption: generated futures faithfully reflect arbitrary valid actions. Existing benchmarks are typically confined to expert demonstrations, leaving off-exp...
Si-Xiang Chen, Jia-Ming Liu, Ji-Xin Wu et al.· 0 citations
Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradig...
A unified generative framework that leverages pre-trained Diffusion Transformer priors to achieve high perceptual quality at extremely low bitrates, achieving state-of-the-art perceptual fidelity with rich spatial details and robust temporal consistency.