Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. Here we show that RLVR on facts increases extraction of personally identifiable information (PII) the instruct model had already memorized. We first confirm that instruct models have already memorized PII but leave them latent, rarely surfacing one when asked. We then apply RL on benign factual data that contains no PII of any kind, and re-probe: a targeted probe over name->email pairs, and an untargeted free-recall prompt that simply asks the model to list the addresses it knows. PII extraction rises sharply under both: on DeepSeek-V3.1, verbatim recall@k increases from 0.155 to 0.370, a 2.4x gain. The effect scales with model size: across three models spanning 8B to 671B parameters, absolute leakage is largest in the biggest model. Meanwhile model's reasoning abilities and refusal rates are retained, indicating that RL selectively changes which memorized information is accessible rather than broadly altering the model. In summary, memorized private data can be made markedly more extractable by training that never touches it. This gives an adversary a route to memorized data that requires no privacy-relevant training signal and no access to the data itself -- only the ability to fine-tune on something innocuous.
This work evaluates SymBuild in three construction domains: computer-aided design (CAD) assembly, Mini-Programs, and exact-fill packing, and test additional framework instantiations in all four domains, demonstrating that SymBuild is an effective, analyzable method for anytime verified construction.
SERUM is the first system to produce interpretable process models from unstructured egocentric screen video without manual annotation, opening a scalable pathway for user modeling and behavioral understanding in the wild.
Andy J. Phu, Karin de Langis, James C Mooney et al.· arXiv.org· 0 citations
A2TTA is proposed, an Anchored-and-Agile Test-Time Adaptation framework for evolving traffic sensor networks, which transforms topology-induced forecasting errors into an expandable output calibration problem and separates tem- poral adaptation into persistent global correction and agile context-specific specialization.
Du Yin, Xiachong Lin, Yuejie Tan et al.· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
It is proved that no estimator computable from the information a deterministic scheme retains is consistent for its own eviction error: evicted values can be altered so that everything retained is unchanged while the true attention-output error grows without bound.
A strong, fixed, rule-based expert is built for Gin Rummy and used only as a yardstick, never for training, and the result is a lightweight, game-agnostic recipe that trains competitive agents without training on the expert, for any game a small model can handle, reported with robust statistics and released as a reusable package.
Nima Kelidari, M. Haghi, Mahdi Salmani· arXiv.org· 0 citations
Bounded unfiltered teacher continuations at learner-induced contexts improve over pure behavioral cloning at matched budgets and suggest that a few teacher steps, placed at learner-induced contexts, can be a more cost-efficient supervision allocation than longer or more heavily curated teacher completions.
Junze Ye, Jiayi Cheng, Miao Lu et al.· arXiv.org· 2 citations
CARE is introduced, a reference-conditioned controller that separates program synthesis from experiment selection and attains the lowest normalized regret, the highest normalized best-so-far AUC, and the highest Top-1% Success@15 among the evaluated methods.
Guanyu Liu, Weiyi Kong, Chao Tang et al.· 1 citation
AlignGAD is proposed, a zero-shot generalized graph anomaly detection framework that aligns heterogeneous node features and normalizes graph signals in the spectral domain and demonstrates the effectiveness of AlignGAD under the zero-shot GAD setting.
Phan Nguyen, Dat Cao, Hien Chu et al.· arXiv.org· 0 citations
This submission documents the divide-and-conquer modeling strategy developed for the CTF-4-Science Lorenz Chaotic Systems Challenge at AI-DEEDS 2026, which shows that bounded, scenario-specific updates can outperform broad model replacement on mixed chaotic forecasting benchmarks.
GRZO is a Group-Relative Zeroth-Order optimizer that draws one pseudo-independent perturbation per mini-batch example and aggregates the per-example losses through group-relative normalization, raising the effective gradient-direction count from one to the batch size at no additional forward cost while preserving inference-level memory.
L. Tan, Yequan Zhao, Yifan Yang et al.· arXiv.org· 0 citations
This work proposes LaRA, a layer-wise representation analysis framework for detecting contamination in RL post-trained LLMs and finds that contamination produces progressive geometric deviations across layers, including amplified perturbation sensitivity, stronger directional collapse, and enhanced local rigidity.
Minju Gwak, Minseok Kwak, Dongseok Lee et al.· arXiv.org· 0 citations