Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories
This work proposes Agentic-DPO, a lightweight offline agent policy optimization method that turns expert trajectories into state-conditioned preference supervision, and introduces Policy-Preserving Augmentation (PPA), which renders the same latent trajectory under multiple schemas while keeping the expert policy fixed.