Exploring Sparse Autoencoders in Text-Based Causal Confounding Adjustment
This work proposes a novel causal adjustment pipeline that iteratively selects a minimal set of SAE features via conditional independence tests, and finds that SAE representations achieve better adjustments than alternative representations in standard semi-synthetic evaluations with binary confounders, and their interp...