This work proposes a novel causal adjustment pipeline that iteratively selects a minimal set of SAE features via conditional independence tests, and finds that SAE representations achieve better adjustments than alternative representations in standard semi-synthetic evaluations with binary confounders, and their interpretability offers opportunities for falsification.
Abstract
In many settings, studying causal questions based on text data requires adjusting for confounding information within texts. Yet there is a tradeoff in constructing text representations for adjustment: they must be sufficiently large and/or dense to preserve the confounding variables necessary for unbiased effect estimation, but sufficiently small and/or sparse to satisfy finite-sample overlap and yield low-variance estimates. To address this tradeoff, we turn to sparse autoencoders (SAEs), and propose a novel causal adjustment pipeline that iteratively selects a minimal set of SAE features via conditional independence tests. We find that SAE representations achieve better adjustments (lower bias and and higher coverage) than alternative representations in standard semi-synthetic evaluations with binary confounders, and their interpretability offers opportunities for falsification. We also introduce a more realistic semi-synthetic evaluation that uses multi-label data as the unobserved confounders and find off-the-shelf adjustment methods require increased investigation for these more complex settings. Code: https://github.com/mianzg/sae-text-confounder
We introduce GenAI-Powered Inference (GPI), a statistical framework for causal inference using unstructured data, including text and images. GPI leverages open-source pretrained Generative AI (GenAI) models-such as large language models and diffusion models-not only to generate unstructured data at scale but also to ex...
Kosuke Imai, Kentaro Nakamura· Proceedings of the National...· 2 citations
Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental samples, making it difficult to estimate heterogeneous effects from high-dimensional covariates. In such settings, policymakers and medical practitioners often succumb to...
Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez et al.· 0 citations
Large language models can generate several records within one autoregressive completion, making earlier answers available as context for later records. This paper shows that such shared-completion batching can distort relationships among variables in the resulting synthetic data, using controlled tests on synthetic sur...
This paper proposes a simple solution to biased estimates of instrumental variables constructed from multiple measures created on independent splits of the original data: instrumental variables constructed from multiple measures created on independent splits of the original data.
Control variables are widely used in statistical modelling to account for omitted variable bias of known confounders. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictio...
Manuel Pfeuffer, R. Rane, Kerstin Ritter et al.· 0 citations
Observational datasets frequently contain many baseline variables, yet investigators estimating causal effects may not know which variables to include in the adjustment set. Confounding information may also be distributed weakly across many variables. Propensity scores can simplify adjustment by reducing high-dimension...
Eric V. Strobl· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.