Skip to content

Author

Johann D. Gaebler

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

From Constitutions to Control: Interpretable Rewards for Aligning Language Models

Current approaches to aligning language models often make it hard to know what behavior is being rewarded or to change that reward in a targeted way. In particular, standard preference-based methods collapse multiple considerations into aggregate human judgments, obscuring what drives the resulting reward, while princi...

Johann D. Gaebler, C. Isley, Max Lamparth et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models

A central concern with language models is sycophancy: their tendency to defer to users'views at the expense of independent substantive judgment. In parallel, work on social sycophancy has focused on behaviors such as validation and positivity that may signal inappropriate deference. Yet the markers of social sycophancy...

C. Isley, Johann D. Gaebler, Max Lamparth et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.