CREST is proposed, an inference-time alignment method that steers base model hidden representations using safety directions extracted from a guidance model of any family, avoiding token-level structural limitations entirely and outperforming baselines by up to 22.2\% on safety benchmarks.
This work proposes Decentralized Barrier Follow-the-Regularized-Leader (Dec-BFTRL), and evaluates each agent's played action against the average of all local objectives, with applications to online continuous diminishing-return (DR) submodular maximization.
Yiyang Lu, M. Pedramfar, Vaneet Aggarwal· 0 citations
This work quantifies the interaction between prosodic features and syntactic representations as their mutual information, and provides a general-purpose framework for estimating this quantity over large speech-text corpora using multimodal language models.
Junghyun Min, Alex Warstadt, Tamar I. Regev et al.· 0 citations
Motus2 is presented, a self-evolving general world model for dexterous manipulation that combines egocentric data scaling and closed-loop general world model scaling to provide a general path toward self-evolving dexterous manipulation.
Hong-Zhe Bi, Zikun Zhou, Yihao Tang et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
VIBE is introduced, a novel text-and-video-to-music (T+V2M) generation model that leverages a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints and soft, subjective qualities with a structured 5-stage training curriculum.
A novel framework that aligns multi-trajectory supervision with policy optimization, and introduces two complementary mechanisms: feasibility-first advantage assignment and dynamic distillation to ensure that expanded trajectory supervision is effectively absorbed during policy optimization.
Tian Zhang, Zhuo Huang, Hong-Rui Ye et al.· 0 citations
This work theoretically shows that directly leveraging PRM score is vulnerable to verifier noise through an extreme-value effect: non-viable prefixes become more likely to receive spuriously high scores as reasoning depth increase, leading to a training-free robust process supervision method that preserves promising alternatives when step-level scores are noisy.
Balance of Benchmarks (BoB) is introduced, which embeds benchmark descriptions and assigns each benchmark an inverse-density semantic weight, providing a principled foundation for task-aware and multiplicity-robust model evaluation.
A multi-solver disagreement reward using a heterogeneous ensemble varying in model capacity and sampling temperature is proposed, which enables the Challenger to discover questions targeting true capability boundaries, producing a curriculum that forces downstream Solvers to develop robust reasoning strategies generalizing across problem types.
This work presents TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks and introduces the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the evaluation objectives.
Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh et al.· 0 citations
The results show that PURGE consistently reduces hallucinations and spurious-correlation-driven errors while maintaining or improving overall performance in most evaluated settings, providing both a reusable evaluation protocol and an effective mitigation framework for more reliable LVLMs.
Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan et al.· 0 citations
A fully open-source pipeline that reproduces the behavior of a cloud-native, event-driven system -- file arrival triggering a message, a message triggering compute -- entirely on commodity hardware, using Apache Kafka and a filesystem-watching poller in place of managed cloud triggers is described.
B. S. Shaikh, M. Mascarenhas, Nuzhat F. Shaikh· 0 citations