Skip to content

Category

artificial intelligence

6,313 papers

#artificial intelligence Preprint Aug 2026

Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering

CREST is proposed, an inference-time alignment method that steers base model hidden representations using safety directions extracted from a guidance model of any family, avoiding token-level structural limitations entirely and outperforming baselines by up to 22.2\% on safety benchmarks.

Jin Gan, Xin Li, Jun Luo · 0 citations
#artificial intelligence Preprint Aug 2026

Dec-BFTRL: Squre-Root Regret for Decentralized Online Upper-Linearizable Optimization under Separation Access with Application to Continuous Submodular Maximization

This work proposes Decentralized Barrier Follow-the-Regularized-Leader (Dec-BFTRL), and evaluates each agent's played action against the average of all local objectives, with applications to online continuous diminishing-return (DR) submodular maximization.

Yiyang Lu, M. Pedramfar, Vaneet Aggarwal · 0 citations
#artificial intelligence Preprint Aug 2026

Using Prosody to Predict Syntactic Structure

This work quantifies the interaction between prosodic features and syntactic representations as their mutual information, and provides a general-purpose framework for estimating this quantity over large speech-text corpora using multimodal language models.

Junghyun Min, Alex Warstadt, Tamar I. Regev et al. · 0 citations
#artificial intelligence Preprint Aug 2026

VIBE: Video Instruction-aligned Background music gEneration

VIBE is introduced, a novel text-and-video-to-music (T+V2M) generation model that leverages a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints and soft, subjective qualities with a structured 5-stage training curriculum.

Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Aligning Multi-Trajectory Supervision with Policy Optimization for VLA Driving

A novel framework that aligns multi-trajectory supervision with policy optimization, and introduces two complementary mechanisms: feasibility-first advantage assignment and dynamic distillation to ensure that expanded trajectory supervision is effectively absorbed during policy optimization.

Tian Zhang, Zhuo Huang, Hong-Rui Ye et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Mitigating Over-Optimization in PRM-Guided Search in Mathematical Reasoning by Optimizing the Guide

This work theoretically shows that directly leveraging PRM score is vulnerable to verifier noise through an extreme-value effect: non-viable prefixes become more likely to receive spuriously high scores as reasoning depth increase, leading to a training-free robust process supervision method that preserves promising alternatives when step-level scores are noisy.

Taejong Joo, Diego Klabjan · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Uncertainty: Multi-Solver Disagreement Rewards for Self-Evolving Reasoning Curricula

A multi-solver disagreement reward using a heterogeneous ensemble varying in model capacity and sampling temperature is proposed, which enables the Challenger to discover questions targeting true capability boundaries, producing a curriculum that forces downstream Solvers to develop robust reasoning strategies generalizing across problem types.

Vinoth Selvendran, Zhan-Ming Zhang · 0 citations
#artificial intelligence Preprint Aug 2026

TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

This work presents TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks and introduces the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the evaluation objectives.

Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models

The results show that PURGE consistently reduces hallucinations and spurious-correlation-driven errors while maintaining or improving overall performance in most evaluated settings, providing both a reusable evaluation protocol and an effective mitigation framework for more reliable LVLMs.

Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.