Skip to content

Author

D. Africa

We have 3 of 32 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Item Response Theory for AI Safety

Overall, it is shown IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which frontier labs and evaluators adopt.

J. Rivera, Neil Shah, D. Africa et al. · 1 citation
Preprint Aug 2026

Measuring Activation Control in Large Language Models

The Activation Controllability Benchmark is introduced to quantify the extent to which models can modulate their residual stream via natural-language instruction, and suggests that control over the activation space itself could become a confound for monitoring as introspective capabilities increase.

Marek Mateusz Kowalski, J. Rivera, Uzay Macar et al. · 0 citations
Jul 2026

Persona Cartography: Charting Language Model Personality Traits in Weight Space

An unsupervised psychometric pipeline is introduced that recovers four interpretable behavioural factors from model rollouts that can be considered in terms of learning, scaling, and composing traits in weight space, providing a bridge between personality measurement, model editing, and safety.

L. Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.