Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bi...
Wan-Jiang Weng, Yong-Liang Wu, Xiaofeng Tan et al.· 0 citations
By combining HRAP with RMR, ScenePilot offers an efficient alternative to one-shot generation and heavy full-scene optimization, improving physical plausibility, functional coherence, and controllability while preserving diversity.
Jiawei Zhang, Hong-Song Wang, Pan Zhou· 0 citations
TransMem is proposed, a lightweight inference-time parametric memory module that transforms sparse historical hidden states from a frozen LLM backbone into reusable memory representations and introduces evidence-conditioned self-distillation to learn transferable memory utilization rather than task-specific knowledge.