Skip to content
Preprint

Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders

Aug 2026 · 0 citations · 19 references
Computer Science

TL;DR

A query-relation deletion can improve the edited slot while reducing the utility of the prediction set that contains it, and selection-conditional deletion sensitivity whose persistence depends on the readout and intervention operator is studied.

Abstract

A query-relation deletion can improve the edited slot while reducing the utility of the prediction set that contains it. We study this tension in two related ResNet-50 DETR-family checkpoints using recorded, selection-conditional evidence from 710 paired image-relation units per checkpoint. The primary comparison subtracts a matched active control, which deletes the same leader source at a different recorded recipient, from the selected target deletion. It is therefore a composite contrast rather than a same-recipient placebo. The target-minus-control contrast is locally positive and fixed-assignment negative in both checkpoints. The opposite-sign pattern occurs within 302/710 DETR units and 460/710 DINO units. After rematching, the corresponding counts are 285/710 and 433/710. Rematching and native selection absorb enough of the mean loss for DETR intervals to cross zero, whereas DINO intervals remain negative, so persistence across readouts differs by checkpoint. A fixed-map comparison between hard deletion and a mass-preserving edit also differs before rematching. That comparison is conditional on the outcome-blind map and does not establish same-dose transport. Local intervention success therefore does not determine the consequence for a jointly decoded set. The supported conclusion is selection-conditional deletion sensitivity whose persistence depends on the readout and intervention operator. We do not identify an intervention-invariant edge mechanism, detector-level degradation, population prevalence, or the value of a training-time regularizer.

View source

Similar papers

#machine learning Preprint Sep 2026

The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits

Hybrid language models can improve capability as well as efficiency, raising the question of how architectural complementarity becomes learned computation. We examine the established induction roles of Carrying predecessor information, Matching a source by content, and Copying its value. How are these position-sensitiv...

Ke Cheng, Xin Xu, Yi-Xiao Chen et al. · 0 citations
Preprint Aug 2026

Loreley: Repository-Scale Program Evolution with Quality-Diversity Search

This work compares configured Loreley QD, sequential champion editing, and independent root proposals in a matched Zstandard experiment and finds that Sequential had the highest observed 48-job mean and median and established a QD advantage.

Mo Chen · 0 citations
#natural language process... Preprint Sep 2026

MeRoTune: RoPE-Safe Merging with a Tunable Dial

When you merge two fine-tuned models from the same base checkpoint by simply averaging their weights, you implicitly assume their attention subspaces are still aligned. Recent work attempts to fix misalignments by learning an invertible correction matrix, $M$, for each model's query and key projections. This correction...

Salman Faroz · 0 citations
#machine learning Preprint Sep 2026

A budget-dependent crossover between coverage- and response-based training-set selection for machine-learned interatomic potentials

Selecting compact training sets for machine-learned interatomic potentials requires deciding whether to preserve structural diversity or target configurations on which models disagree. The better choice can depend on how much data is retained, making a comparison at one training-set size insufficient. Here we link sele...

Jiancheng Bi, Alin M. Elena · 0 citations
#machine learning Preprint Sep 2026

LARC: Low-Rank Adaptive Residual Connections for Learning in Frozen Models

Low-Rank Adaptive Residual Connections (LARC) give a frozen model a compact numerical state that can learn from feedback. The map $h+BAh$ adds a low-rank correction to a hidden representation. A slow state $\rho$ learns starting factors across tasks; a private fast state $\Phi$ copies them, changes with feedback, and r...

Jun-Yi Zou, Avrova Donz · 0 citations
Preprint Aug 2026

Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication

Generative visual-token communication reduces transmission load by sending only selected discrete tokens and reconstructing missing content at the receiver. However, existing token-selection criteria based on local uncertainty, importance, or diversity do not directly determine whether changing the current selection im...

Jia Guo, Xiaohan Zhao, Changwang Liu et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.