HDR-RoPE is proposed, which extends RoPE from independent 2D rotations to higher-dimensional rotations and introduces a Paley-I orthogonal basis to obtain balanced, isotropic, and dense phase mixing within each rotation subspace and significantly enhances channel coupling and rotational degrees of freedom while maintaining orthogonal stability and the relative position closure property.
Yixing Li, Ruobing Xie, Yu-Dong Zhang et al.· 0 citations
LLMODE is proposed, a token-efficient framework for irregular spatio-temporal forecasting with a frozen LLM backbone that shows competitive overall performance, with clearer advantages under sparse or dynamically complex irregular sampling.
Di Zhang, Jing-Yang Zhang, Zi-Qian Wang et al.· 0 citations
The construction establishes that learning can change effective epistemic reach even when primitive affordances and deployment resources are held fixed, and opens a complementary evaluation question for learning systems: not only what they infer from available evidence, but what informative evidence experience teaches them to bring within reach.
Professional basketball is the case study, chosen for its data rather than the league, and five public sources are fuse into one per-shot dataset of 4.23M shots over 21 seasons, finding that the analytics tools of professional teams stay out of reach.
Yi-Bo Gong, Congyu Guo, Jiachen Ding· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
It is shown that continual unlearning operations accumulate geometric constraints in parameter space, leading to saturated subspaces that restrict future updates, a critical barrier to the long-term reliability of machine unlearning systems and motivate the development of plasticity-preserving unlearning algorithms.
Ying-Dan Shi, Xiang-Dong Xu, Kaize Ding et al.· 0 citations
The proposed update incorporates the objective gradient inside the denoising step, yielding an inference-time method that uses only a pretrained denoiser and gradient evaluations and is analyzed as an inexact projected-gradient method for constrained optimization over learned feasible geometries.
R. Zhang, Jiawei Zhang, Gioele Zardini et al.· 0 citations
It is proved that the learning dynamics and the distillation error $\Ets$ are exactly invariant to $\dmiss$, whereas the true error $\Etzs$ and the gap $\Delta=\Etzs-\Ets$ are strictly increasing in $\dmiss$, with a rate that is amplified linearly by the complexity $M_0$ of the true teacher.
Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluations that frontier-model governance depends on. The Elicitation Game found that fine-tuning elicits hidden capability from sandbagging model organisms whereas additive activation steering fails. We revisit that verdict with reference-grafting, which sets an activation's coordinate along a contrast direction to the value it takes in an honest reference, at a small set of circuits chosen by active learning. Across eleven fine-tuned password-locked models (1.5-32B, three architecture lineages), it recovers +94 to +101% of the honest-sandbagging gap, matching fine-tuning elicitation without weight updates or training labels; two to five paired examples suffice to fit the direction. Similar recovery holds for reinforcement-learning-induced sandbagging and for password-locked code generation. Grafting works because the fine-tuned lock is a thresholded gate: held-out accuracy stays at the sandbagged level until the grafted coordinate crosses a threshold near the honest reference, which is why additive steering and zeroing the coordinate often fail. The direction tracks how the sandbagging was induced rather than what is withheld -- fit on grade-school science (ARC) it elicits withheld hazardous knowledge (WMDP), yet prompting, fine-tuning, and reinforcement learning each carry a different direction. Circuit-breaking marks the boundary: it reroutes activations on every forward pass, so the fixed edits we test are re-broken downstream and do not restore coherent generation.
Linh Le, Hong-Fu Tan, David Williams-King· 0 citations
Method is introduced, which augments SOAP-style preconditioning with a scalar secant-energy correction adapted to Kronecker geometry and an adaptive basis update followed by variance-state downscaling, and is positioned as a scalable option for stiff, high-accuracy physics-informed training, rather than a uniform replacement for existing optimizers.
Guang-Yuan Wang, Mads Toftrup, Sebastian Loeschcke et al.· 0 citations
Geometry finally makes a commanded 3D target a natural goal interface: it is constructed the goal latent from the target and the current latent, at no cost in success rate, without a goal observation.
F. F. Oberweger, Michael Schwingshackl· 0 citations
It is argued that LASSO, not the highest-discriminating model, is the model best suited to direct clinical deployment, and lessons for the machine learning and healthcare community regarding data infrastructure, model selection, and value of calibration and interpretability in high-stakes decision support are presented.
Asra Aslam, Volodymyr Chapman, M. O'Connell et al.· 0 citations
This work proposes a fully distributed continuous-time algorithm for shared linear equality constraints that converges without multiplier exchange and reaches any GNE, reducing communication overhead and improving privacy.
Sho-An Yin, Mingyi Hong, Nicola Elia· American Control Conference· 1 citation