What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation
When a language model explains an answer it has already given, does it reuse the computation that produced the answer or reconstruct a story from the answer alone? Attribution, transportability and recoverability are each compatible with causal use without establishing it. We propose an evidence standard: pair each pos...