Skip to content
Review

Dynamic Discrete Choice and Inverse Reinforcement Learning: Inferring Preferences and Beliefs From Human Behavior

Aug 2026 · 0 citations · 106 references
Economics

TL;DR

It is shown that the (soft Q-learning) framework now prevalent in IRL is closely related to DDC models under additive extreme value preference shocks, yielding the same softmax (multinomial logit) choice probabilities and smooth Bellman equations that underpin structural estimation in economics.

Abstract

This article surveys two deeply connected literatures that approach the same fundamental problem from different disciplinary traditions: dynamic discrete choice (DDC) in structural econometrics and inverse reinforcement learning (IRL) in machine learning. Both seek to infer the preferences of decision makers from observed sequential behavior, assuming that individuals act to maximize an expected reward function within a dynamic, uncertain environment formalized as a Markov decision process (MDP). Despite independent origins, the two fields have converged on similar mathematical formulations. We show that the (soft Q-learning) framework now prevalent in IRL is closely related to DDC models under additive extreme value preference shocks, yielding the same softmax (multinomial logit) choice probabilities and smooth Bellman equations that underpin structural estimation in economics. We compare the estimation and computational methods developed in each field. DDC has emphasized maximum likelihood estimation, conditional choice probability estimators, and policy iteration methods. IRL has developed scalable alternatives, including maximum entropy methods, adversarial approaches, and model-free temporal difference estimators that extend to high-dimensional state spaces using deep neural networks. Model-free IRL estimators that combine temporal difference learning with classical two-step methods from econometrics represent a promising direction for bridging the two literatures. Both fields confront shared foundational challenges: the identification problem, whereby multiple reward functions can rationalize the same observed behavior, and the curse of dimensionality in solving the underlying MDP. We believe that cross-fertilization offers substantial opportunities for methodological progress in both fields.

View source

Similar papers

Preprint Aug 2026

Neural-Bayesian Structure Learning for Discrete Choice Modeling

Conventional discrete choice and machine learning models are estimated primarily from observational data and typically treat explanatory covariates as parallel inputs, providing no internal mechanism for determining how related attributes should adjust when one is deliberately changed. This paper proposes Neural-Bayesi...

Hyun-Ah Yun, Eun Hak Lee, Jia-Ru Zhang et al. · 0 citations
Review Aug 2026

Learning Sequential Mobility Choice: A Review of Route and Activity Choice through Inverse Reinforcement Learning and Imitation Learning

Route and activity choice are distinct transportation problems that both require models of feasible decisions unfolding over networks and time. This critical integrative review connects transportation choice modeling with inverse reinforcement learning (IRL) and imitation learning (IL), while distinguishing evidence fr...

H. Tran, Viet The Bui, Tien Mai · 0 citations
Preprint Aug 2026

Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes

A nonparametric distributional Bellman optimality operator for JMDPs is defined, and it is proved that when the induced marginal MDP has a unique optimal policy, its iterates converge in Wasserstein distance to the optimal joint return law.

Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi et al. · 0 citations
#machine learning Preprint Aug 2026

Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View

The empirical gap between these method families is identified as a variance-reduction effect rather than a difference in RL principle, and a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones...

Yi-Xian Xu, Yuanrui Zhang, Sheng-Jie Luo et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.