Inferring the evolving objectives that organize animal behavior
Abstract
Understanding how the brain supports flexible, goal-directed behavior is a central aim of neuroscience. Yet such behavior is highly variable and unfolds over long timescales, making it challenging to understand its organization. Advances in behavioral tracking quantify movements but do not reveal the objectives that guide them. Motivated by this, we develop inverse reinforcement learning (IRL) into a general framework for inferring these objectives from naturalistic behavior. Our approach represents an animal’s objectives as internal rewards, and infers these rewards as well as the policies they induce from observed behavior. It applies across diverse behavioral settings, from freely moving behaviors to structured tasks. It can infer objectives that switch over time or gradually change with experience. In mice capturing crickets, it identified a small set of objectives that the animals switched between over the course of a trial. In mice learning a cued labyrinth, it inferred how the animals’ internal rewards changed as learning progressed. Across both datasets, the models predicted held-out actions and generated extended behavior that reproduced key features of animal performance. By transforming long, variable behavioral sequences into time-resolved estimates of the animal’s objectives, our approach provides a quantitative basis for studying how the brain organizes natural behavior.