An Empirical Study of Counterfactual Self-Explanations in LLMs
This work evaluates ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales and shows that model scale is the strongest determinant of explanation quality.
Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis-Mastromichalakis et al.
· 0 citations