The Activation Controllability Benchmark is introduced to quantify the extent to which models can modulate their residual stream via natural-language instruction, and suggests that control over the activation space itself could become a confound for monitoring as introspective capabilities increase.
Marek Mateusz Kowalski, J. Rivera, Uzay Macar et al.· 0 citations
An unsupervised psychometric pipeline is introduced that recovers four interpretable behavioural factors from model rollouts that can be considered in terms of learning, scaling, and composing traits in weight space, providing a bridge between personality measurement, model editing, and safety.
L. Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.