Skip to content
Preprint

Double Machine Learning with High-dimensional Interactive Fixed Effects

Aug 2026 · 0 citations · 55 references
Economics

TL;DR

The method combines projection-based defactorisation of the data, in the spirit of Common Correlated Effects, with a Neyman-orthogonal score function and cross-fitting procedure, and accommodates low-rank factor structures in outcomes and treatments alongside high-dimensional, potentially nonlinear covariate effects estimated by machine learning algorithms.

Abstract

Factor structures are central to empirical work in economics and finance, and are usually used to model time-varying unobserved heterogeneity through interactive fixed effects (IFE). Existing IFE estimators rest on low-dimensional and linear specifications in the covariates, assumptions which are increasingly restrictive in applications drawing on rich datasets with controls of unknown functional form. This paper develops a Double Machine Learning estimator for the high-dimensional partially linear panel model with interactive fixed effects (panel DML-IFE). The method combines projection-based defactorisation of the data, in the spirit of Common Correlated Effects (CCE), with a Neyman-orthogonal score function and cross-fitting procedure, and accommodates low-rank factor structures in outcomes and treatments alongside high-dimensional, potentially nonlinear covariate effects estimated by machine learning algorithms. Monte Carlo simulations show that panel DML-IFE outperforms conventional IFE estimator outside the correctly-specified linear case, with bias reduction driven primarily by the time and covariate dimensions. An empirical application to U.S. stock returns shows that several effects documented under linear specifications lose statistical significance once high-dimensional nonlinear confounding and the presence of IFE are jointly accounted for.

View source

Similar papers

Preprint Aug 2026

High-Dimensional Panel Data Models with Interactive Fixed Effects: Beyond the Linear Case

Modern economic panel data sets are often high-dimensional: they contain information on a wide variety of control variables whose number may even exceed the sample size. Nevertheless, the literature on econometric methods for high-dimensional panels is quite limited. In this paper, we study high-dimensional panel model...

Maximilian Ruecker, M. Vogt, Oliver B. Linton · 0 citations
Preprint Sep 2026

Covariate Selection for Doubly Robust Double/debiased Machine Learning Estimators for Causal Inference

High-dimensional data create challenges for causal effect estimation because identifying the covariates needed for correct model specification becomes increasingly difficult. Double/debiased machine learning (DML) facilitates the use of machine learning (ML) for causal inference by mitigating regularization and overfit...

Muwon Kwon, Peter M. Steiner · 0 citations
Preprint Sep 2026

Design-Assisted Regression

We consider regression problems in which the marginal distribution of the covariates is informative for estimation and variable selection, rather than merely auxiliary. Motivated by random-design, high-dimensional, and latent-effect settings, we propose a general design-assisted regression framework in which the estima...

S. Ye, Guan-Bo Wang, Cong Zhang et al. · 0 citations
Preprint Sep 2026

Orthogonal Moments in Likelihood Models

Many models, such as fixed-effect models for panel or network data, are hard to estimate because they feature nuisance parameters that are both numerous and estimated imprecisely. This, in general, causes an incidental-parameter problem in the estimator of the parameters of interest. The problem can be alleviated by wo...

St'ephane Bonhomme, Koen Jochmans, M. Weidner · 0 citations
Nov 2026

Debiased estimation and variable selection under function-on-scalar linear regression models with ultrahigh-dimensional covariates subject to measurement error

In real-world applications, data are often error-contaminated; naively applying conventional methods without accommodating the measurement error effects often yields inconsistent estimates. Biased results can be further exacerbated by the ultrahigh-dimensionality of covariates. Focusing on the widely used function-on-s...

Yifan Sun, Grace Y. Yi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.