Physical interaction quality is central to deformable-object manipulation, yet most benchmarks evaluate task success alone. A policy may complete the task while allowing slip or causing excessive compression. A primary bottleneck is the absence of visuo-tactile datasets that pair policy-visible contact observations wit...
Bowen Jing, Ming-Xin Wang, Ruiyang Hao et al.· 0 citations
Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generati...
Qi-Ze Yu, Lian-Rui Fan, Bo-Yu Chen et al.· 0 citations
SCALE (Selective Control of Adaptation via Local Entropy), an entropy-guided adaptation-strength-control method that freezes the pretrained model and the SFT delta and learns bounded token- and module-specific gates by minimizing predictive entropy alone is proposed.
Cun-Chun Li, Hao-Nan He, Yi-Fan Gao et al.· 0 citations
UMI-Bridge, which uses UMI as an intermediate domain to align representations according to action equivalence rather than pixel similarity, supports action-anchored latent alignment for data-efficient robot learning and UMI-to-robot task transfer.
Hai-Yi Liu, Jin-Ming Ma, Ke Rui et al.· 0 citations
JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). Strong decoding of task variables, however, does not guarantee that this particular cost ranks candidate action sequences by real task progress. We call the latter property \emph{decision-metric al...
Jia-Wei Wang, Ke Rui, Yu-Shen Zuo et al.· 4 citations
HeteroGenManip is proposed, a task-conditioned, two-stage framework designed to decouple initial grasp from complex interaction execution, and achieves robust intra-category shape and pose generalization.
GaussianWAM is proposed, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field and improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation.
Zi-Jian Zhang, Yu-Qing Jiang, Wei-Tao Zhou et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.