Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during Knowledge Distillation (KD) from teacher to student remains largely unexplored. Thus, we fill th...
Anjila Budathoki, Manish Dhakal, Benjamin M. Ampel et al.· 0 citations
The experiments show that authority bias is substantial and directional, varies markedly across models, and is only partially addressable through prompt-level debiasing, and document a say-do gap.
Uthman Jinadu, Parsa Ghazvinian, Anjila Budathoki et al.· 0 citations
It is suggested that predictive performance, attribution plausibility, and mechanistic faithfulness characterize different aspects of model behavior and should be evaluated separately when studying explainability in media bias detection.
Tinghao Chen, Raina Zhang, Benjamin M. Ampel et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.