The OISD framework is proposed, which improves reasoning by transferring on-policy predictive signals from the final layer to intermediate representations and employs signed advantage-weighted Jensen--Shannon alignment to distill informative intermediate representations while preserving policy consistency under a unified acting policy.
Xin-Yu Liu, Darryl C. Jacob, Yang Zhou et al.· arXiv.org· 0 citations
It is shown that inter-layer redundancy can be either localized or globally distributed depending on the LLM architecture, and Representation Locality Score (RLS) is introduced, derived from global inter-layer hidden-state similarity.
Vincent-Daniel Yun, Youngrae Kim, Woosang Lim et al.· arXiv.org· 1 citation· ⚡1
Student-Centric Answer Sampling (SCAS) is proposed, a framework that selects from verified teacher-generated answers according to their estimated student-centric learning cost and is derived by a token-wise gradient decomposition and used to guide answer selection during training.
Zhengyu Hu, Zheyuan Xiao, Linxin Song et al.· 0 citations
ProFIL (**Pro**be-**Filtered Reinforcement Learning) is introduced to reduce theater, increase chain-of-thought faithfulness, and shrink chain length in a single, drop-in extension to Group Relative Policy Optimization (GRPO).
Swapnil Parekh, Naman Goyal· arXiv.org· 2 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
AOPD replaces ineffective negative reinforcement with localized divergence minimization in non-positive advantage regions while preserving positive reinforcement learning and maintains higher policy entropy during training and better capability retention during sequential tool-use adaptation.
Nan Jia, Haojin Yang, Xing-Chen Ma et al.· arXiv.org· 18 citations· ⚡5
The results suggest transformers rotate semantic content into spectrally quiet regions during contextualized processing, where, in some architectures, interventions may reduce grammatical disruption relative to high-variance steering.
Pratyush Acharya, Nuraj Rimal, H. Dhakal· 0 citations
AirFM-DDA is proposed, an Air-interface Foundation Model in the Delay-Doppler-Angle (DDA) domain, which reparameterizes CSI into the DDA domain to resolve multipath components along physically meaningful axes and employs window-based attention with frame-structure-aware positional encoding.
Kejia Bian, Meixia Tao, Jianhua Mo et al.· arXiv.org· 6 citations· ⚡1
This work pairs the brain emulator with large language models that generate news headlines from linguistic parameters such as valence, arousal, and dominance and shows that these parameters can be recovered from predicted brain maps, demonstrating that the emulator's synthetic neural encodings preserve information about the controlled stimulus dimensions.
Niels Bracher, Xavier Intes, Stefan T. Radev· arXiv.org· 0 citations
This work describes coordinated perturbation as a budgeted subset-selection problem over pairwise observations and introduces an Adaptive Subset Selection Attack (ASSA) as a scalable search heuristic for probing high-impact perturbation sets.
These results suggest that in diversity-aware multi-armed bandits, e.g., for generative model selection, exploration can arise intrinsically from the objective's geometry, particularly for widely used metrics such as FID and Vendi where tight confidence bounds are difficult to construct.
This work presents SELA, a neuro-symbolic VLM agent framework that iteratively grounds primitives from signal visualizations and composes them under ELT constraints, producing both event intervals and faithful tree-structured explanations.
Sky Chenwei Wan, T. Hou, Yifei Wang et al.· arXiv.org· 0 citations
Personalized GRPO is introduced, a novel alignment framework that decouples advantage estimation from immediate batch statistics and achieves faster convergence and higher rewards than standard GRPO, thereby enhancing its ability to recover and align with heterogeneous preference signals.
Jialu Wang, Heinrich Peters, A. Butt et al.· arXiv.org· 1 citation