Reusable collider representations can be evaluated through downstream discrimination and probes of retained information, but neither quantity directly tests the behaviour of score templates in a profiled likelihood. We test a specific prediction in a controlled two-channel routing protocol: if reduced physics-label readability in a nuisance branch indicates a more inference-robust representation, it should accompany a smaller profiled signal-strength bias under fixed unmodelled shifts. In a public Compact Muon Solenoid $H\rightarrow ZZ\rightarrow4\ell$ workflow, a downstream split of fixed EveNet embeddings preserves signal/background area under the receiver operating characteristic curve ($0.9894\pm0.0004$) while reducing nuisance-branch physics readability from $0.961\pm0.013$ to $0.593\pm0.030$. Probe-sensitivity and effective-rank controls exclude a failed readout and branch collapse. In a separate top quark jet-tagging workflow, the leakage reduction recurs with preserved task performance. Across two development event shards, however, its Spearman association with maximum absolute profiled bias is $0.036$, and three of six material leakage-improving transitions do not reduce that bias. A one-shot preregistered confirmation on an independently accessed shard produces material leakage reductions in all three paired seeds, while the maximum absolute bias increases in two. Thus, within the tested protocol, latent readability is a useful routing diagnostic but not a likelihood-robustness certificate. The result supports a practical validation rule: claims about inference robustness require a prespecified likelihood-facing stress test and held-out confirmation.
Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated text. We propose a reference-based method that audits bias in hidden-state representations across related model variants, for example before...
M. Jeliński, Jan Dubinski, Maciej Chrabaszcz et al.· 0 citations
Searches for new physics in low-background experiments infer a non-negative signal strength from few events and often report an upper limit. Nominal frequentist coverage requires both a valid interval construction and an adequate data model. We study how controlled model departures affect lower- and upper-endpoint cove...
This work introduces a diagnostic protocol using a minimal, target-label-free additive correction, showing that many apparent zero-shot reasoning deficits are expression failures masking intact internal logic, urging a narrower interpretation of benchmark evaluations.
It is shown the natural way to do this does not work, specify one that survives measurement, then finds that the correction making it work carries more variance than the null it is tested against, and that the correction making it work carries more variance than the null it is tested against.
This work studies 24 open-weight monitors spanning nine pretraining lineages and a 29x range of detection skill, and reports that across six attacker models the gain result holds in all six, the agreement and cancellation results in five of six.
When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experime...
Girish A. Koushik, Diptesh Kanojia, Helen Treharne· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.