An episode-level protocol that scores whether, when, and what a delegate contributes around the participant's actual idea units is introduced, to evaluate online delegation.
Muneeb Khan, F. Kirstein, T. Ruas et al.· 0 citations
This work presents the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry, and introduces three novel metrics that isolate social deduction, reasoning, and deceptive consistency.
Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa et al.· arXiv.org· 0 citations
It is found that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both, and gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging.
Lars Benedikt Kaesberg, Tianyu Yang, F. Wunderlich et al.· 0 citations
The findings suggest that activation steering may provide a practical, low- cost mechanism for extending English-derived safety signals to other languages, and introduce a multilingual translation-and-evaluation pipeline to facilitate future work on cross-lingual safety interventions.
Emma V. Stein, Dominik Meier, T. Ruas et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.