The results indicate how trustworthy LLM-generated explanations are in model-free settings, where the same LLMs are used but no oracle exists to verify them.
The TSExplorer tool enables users to inspect high-dimensional datasets through multiple complementary 2D visualizations derived from high-dimensional feature representations to support a wide range of workflows.
Einari Vaaras, Manu Airaksinen, O. Räsänen· 0 citations
MineAmongUs is introduced, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action, and ARIA is proposed, a configurable VLM-agent harness that exposes five cognitive-component ablation axes and opens a new path for embodied VLM-agent alignment research.
Jaewoo Ahn, Junseo Kim, Hyunseo Kim et al.· 0 citations
This work proposes a decompose-before-reconstruct approach, which significantly improves mesh fidelity and novel-view synthesis and novel-view synthesis, while supporting object-wise modifiability and interactivity.
Minhas Kamal, Hiranya Garbha Kumar, Mahedi Kamal et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
CREST is proposed, an inference-time alignment method that steers base model hidden representations using safety directions extracted from a guidance model of any family, avoiding token-level structural limitations entirely and outperforming baselines by up to 22.2\% on safety benchmarks.
This work proposes Decentralized Barrier Follow-the-Regularized-Leader (Dec-BFTRL), and evaluates each agent's played action against the average of all local objectives, with applications to online continuous diminishing-return (DR) submodular maximization.
Yiyang Lu, M. Pedramfar, Vaneet Aggarwal· 0 citations
This work quantifies the interaction between prosodic features and syntactic representations as their mutual information, and provides a general-purpose framework for estimating this quantity over large speech-text corpora using multimodal language models.
Junghyun Min, Alex Warstadt, Tamar I. Regev et al.· 0 citations
Motus2 is presented, a self-evolving general world model for dexterous manipulation that combines egocentric data scaling and closed-loop general world model scaling to provide a general path toward self-evolving dexterous manipulation.
Hong-Zhe Bi, Zikun Zhou, Yihao Tang et al.· 0 citations
VIBE is introduced, a novel text-and-video-to-music (T+V2M) generation model that leverages a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints and soft, subjective qualities with a structured 5-stage training curriculum.
A novel framework that aligns multi-trajectory supervision with policy optimization, and introduces two complementary mechanisms: feasibility-first advantage assignment and dynamic distillation to ensure that expanded trajectory supervision is effectively absorbed during policy optimization.
Tian Zhang, Zhuo Huang, Hong-Rui Ye et al.· 0 citations
This work theoretically shows that directly leveraging PRM score is vulnerable to verifier noise through an extreme-value effect: non-viable prefixes become more likely to receive spuriously high scores as reasoning depth increase, leading to a training-free robust process supervision method that preserves promising alternatives when step-level scores are noisy.
Balance of Benchmarks (BoB) is introduced, which embeds benchmark descriptions and assigns each benchmark an inverse-density semantic weight, providing a principled foundation for task-aware and multiplicity-robust model evaluation.