Jul 2026
Training Large Language Models for Self-Explanation Faithfulness
It is shown that models can be trained to implicitly identify influential factors and disclose them, offering a scalable path toward reducing unfaithful reasoning in LLMs.
Y. Cheah, María Pérez-Ortiz, N. Siegel et al.
· arXiv.org · 0 citations