Reinforcement learning (RL) trains representations on data selected by the agent's policy, which then uses the resulting returns to guide its next choices. We show that this loop can sustain a lower-return policy even when representation fitting is globally optimal on those data. In a self-confirming superposition trap...
Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident than their corresponding pretrained language models (PLMs). Because task-specific labeled calibration data can be costly or unavailable, the corresponding pretrained PLM...
Linhan Luo, Lequan Lin, Dai Shi et al.· 0 citations
Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are structurally incapable of such jumps, while recent studies challenge both its mechanism and em...
Dai Shi, Xiao-Yu Li, José Miguel Hernández-Lobato· 0 citations
This work presents an exposition of the OSQ problem by summarizing its various formulations in the current literature and categorizing existing solutions into three different types, and summarizes the empirical methods proposed by existing works to verify the efficiency of OSQ mitigation approaches.
Dai Shi, Andi Han, Lequan Lin et al.· IEEE Transactions on Pattern...· 0 citations
Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons and motivates methods for recovering interpretable features from neural activations. Theoretical models typically start with a given set of input features and assumptio...
A formal account of the jump is developed in four steps and measured, proving that jump instances are well-posed and establish a family theorem that certifies instances of unbounded difficulty without enumeration and further formalize when a jump is correct and how successive jumps compound.
Dai Shi, Xiao-Yu Li, José Miguel Hernández-Lobato· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.