PreferenceFlow: Test-Time Guidance of Flow-Matching Robot Policies from Human Interventions
Flow-matching policies can represent complex robot behaviors but remain susceptible to local errors under distribution shift at deployment. Many reinforcement learning approaches to policy improvement require reward signals that are difficult to specify or obtain in real-world manipulation. We present PreferenceFlow, a...