VLAs combine pretrained vision-language representations with action generation to enable language-guided control across diverse tasks, becoming a mainstream paradigm in embodied intelligence. However, multiple studies have reported VLA's substantial declines in task success under camera shifts, revealing a key vulnerab...
Tian-Hang Pan, Xuan Wang, Yi-Wen Pang et al.· 0 citations
On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-toke...
Zhe-Xu Wang, Mao-Lin Luo, Yan-Kun Hong et al.· 0 citations
Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes and temporal evidence. We propose {\ours}, which couples labe...
Ling-Yu Shen, Wei Tang, Fakhri Karray et al.· 0 citations
Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complemen...
Kai Gan, Zi-Hao Zhou, Bo-Tao Ye et al.· 0 citations
Recent advances in multimodal large reasoning models (MLRMs) have demonstrated impressive capabilities on complex multimodal tasks, yet their reliance on long Chain-of-Thoughts (CoTs) often leads to redundant reasoning and high computational cost. Existing chain-based distillation and refinement approaches alleviate re...
Yizhi Wang, Li-Nan Yue, Deng-Bao Wang et al.· Proceedings of the 32nd ACM...· 0 citations
This work introduces MRCL, a Multimodal Reasoning Continual Learning benchmark, and proposes Continual Policy Optimization (CPO), a replay-free framework grounded in a prior-task behavioral KL objective that consistently reduces forgetting while preserving, and in some cases improving, pretrained model capabilities.
This work proposes ComRank, a ranking loss framework for MLCLL, which encourages complementary labels to be ranked lower than non-complementary ones, thereby modeling pairwise label relationships and ensures Bayes consistency under both uniform and biased cases.
Jin Zhu, Yi Gao, Miao Xu et al.· Neural Information Processin...· 0 citations
This work introduces a collective evidence-threshold backdoor paradigm for MAS and Boundary-Conditioned Backdoor Injection, which constructs counterfactual boundary pairs to separate benign behavior before the threshold from the adversarial objective after it, and learns latent progression aligned with evidence.
Jiahao Xiao, Lei Feng, Min-Ling Zhang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.