Skip to content

Exploring Sparse Autoencoders in Text-Based Causal Confounding Adjustment

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

This work proposes a novel causal adjustment pipeline that iteratively selects a minimal set of SAE features via conditional independence tests, and finds that SAE representations achieve better adjustments than alternative representations in standard semi-synthetic evaluations with binary confounders, and their interpretability offers opportunities for falsification.

Abstract

In many settings, studying causal questions based on text data requires adjusting for confounding information within texts. Yet there is a tradeoff in constructing text representations for adjustment: they must be sufficiently large and/or dense to preserve the confounding variables necessary for unbiased effect estimation, but sufficiently small and/or sparse to satisfy finite-sample overlap and yield low-variance estimates. To address this tradeoff, we turn to sparse autoencoders (SAEs), and propose a novel causal adjustment pipeline that iteratively selects a minimal set of SAE features via conditional independence tests. We find that SAE representations achieve better adjustments (lower bias and and higher coverage) than alternative representations in standard semi-synthetic evaluations with binary confounders, and their interpretability offers opportunities for falsification. We also introduce a more realistic semi-synthetic evaluation that uses multi-label data as the unobserved confounders and find off-the-shelf adjustment methods require increased investigation for these more complex settings. Code: https://github.com/mianzg/sae-text-confounder

View source

Similar papers

Open access Sep 2026

Leveraging generative AI for causal inference with unstructured data.

We introduce GenAI-Powered Inference (GPI), a statistical framework for causal inference using unstructured data, including text and images. GPI leverages open-source pretrained Generative AI (GenAI) models-such as large language models and diffusion models-not only to generate unstructured data at scale but also to ex...

Kosuke Imai, Kentaro Nakamura · 2 citations
#machine learning Preprint Sep 2026

Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes

Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental samples, making it difficult to estimate heterogeneous effects from high-dimensional covariates. In such settings, policymakers and medical practitioners often succumb to...

Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez et al. · 0 citations
#machine learning Review Sep 2026

Shared Autoregressive Context Can Distort Relationships in Synthetic Data

Large language models can generate several records within one autoregressive completion, making earlier answers available as context for later records. This paper shows that such shared-completion batching can distort relationships among variables in the resulting synthetic data, using controlled tests on synthetic sur...

Thomas S. Robinson · 0 citations
Preprint Aug 2026

When Predictions Become Regressors: A Split-Sample Correction for Biases in Downstream Inference

This paper proposes a simple solution to biased estimates of instrumental variables constructed from multiple measures created on independent splits of the original data: instrumental variables constructed from multiple measures created on independent splits of the original data.

N. Canen, Ted Enamorado · 0 citations
Preprint Aug 2026

Controlling for Omitted Variable Bias in Deep Neural Networks

Control variables are widely used in statistical modelling to account for omitted variable bias of known confounders. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictio...

Manuel Pfeuffer, R. Rane, Kerstin Ritter et al. · 0 citations
Preprint Aug 2026

COMPACT: Spectral Adjustment Scores from a Complete and Irreducible Causal Criterion

Observational datasets frequently contain many baseline variables, yet investigators estimating causal effects may not know which variables to include in the adjustment set. Confounding information may also be distributed weakly across many variables. Propensity scores can simplify adjustment by reducing high-dimension...

Eric V. Strobl · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.