Skip to content
Preprint

When Predictions Become Regressors: A Split-Sample Correction for Biases in Downstream Inference

Aug 2026 · 0 citations · 30 references
Economics Mathematics

TL;DR

This paper proposes a simple solution to biased estimates of instrumental variables constructed from multiple measures created on independent splits of the original data: instrumental variables constructed from multiple measures created on independent splits of the original data.

Abstract

Prediction-based methods, including Large Language Models (LLMs) and other machine learning techniques, are often used to construct measures of political phenomena that are difficult to quantify directly, such as policy positions in manifestos or emotions expressed on social media. In many applications, these prediction-generated measures are used as explanatory variables in regression models, even though they are measured with error. This leads to biased estimates. In this paper, we propose a simple solution to these biases: instrumental variables constructed from multiple measures created on independent splits of the original data. This approach is theoretically valid, easy to implement, and does not require new data. Through simulations, we show that this approach recovers estimates close to the true values, even in relatively small samples, while the standard approach can produce substantial bias in practice. We illustrate the method by revisiting two applications: whether gendered speech affects legislative outcomes in the German Parliament, and whether political risk influences poverty alleviation programs in China.

View source

Similar papers

Review Jul 2026

Analyzing and Correcting Benevolence Bias in Large Language Models

Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.

Yuanzi Li, Jun-Hao Wang, Minghui Liu et al. · 0 citations
Preprint Aug 2026

Controlling for Omitted Variable Bias in Deep Neural Networks

Control variables are widely used in statistical modelling to account for omitted variable bias of known confounders. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictio...

Manuel Pfeuffer, R. Rane, Kerstin Ritter et al. · 0 citations
Preprint Jul 2026

Lucky or Good? Outcome Noise, Effective Sample Size, and the Attribution of Skill

When do outcome records carry enough signal to support reliable inferences about skill? When they do not, what should evaluators substitute? The framework answering the first question characterizes any decision domain with two parameters: the noise reflected in each outcome and the effective number of independent outco...

Karl T. Ulrich · 0 citations
Jun 2026

Using LLMs for Explainable, Data-Driven Insight Generation from Time Series

Results show that generated explanations approached analyst-written explanations in terms of readability, consistency and persuasiveness, demonstrating that grounded explanation generation for time series forecasting can be achieved at scale without domain-specific fine-tuning.

Ria Mundhra, G. S. dos Santos, Michael Benedikt · 0 citations
Preprint Aug 2026

Counterfactual Analysis via Large Language Models

This paper investigates the application of large language models (LLMs), specifically the GPT-3.5 model, for counterfactual analysis in the online lending context, and utilizes GPT to generate counterfactual ROIs under a set of alternative interest rates.

Zong-Fan Yang · 0 citations
Jul 2026

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

This work introduces OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth.

Seonglae Cho, A. Koshiyama · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.