Uncertainty-Calibrated Hamilton-Jacobi Value Learning for Safe Crowd Navigation
Abstract
Safe navigation in human crowds requires efficiency and reliability. Learning-based policies run fast but can fail under distribution shift, whereas Hamilton–Jacobi (HJ) reachability offers principled safety guarantees but is computationally expensive. We propose a prediction-conditioned reachability framework that bridges these two regimes by learning safety guidance and combining it with a receding-horizon Model Predictive Control (MPC). We adopt pedestrian trajectory prediction for generating future trajectories, which are converted into a time-varying failure geometry (a signed-distance field sequence). Using these, we build a residual value model based on U-Net architecture, which is trained with offline computed HJ reachability dataset, to approximate the safety. To remain conservative under approximation error, the model also outputs a heteroscedastic scale and is conformally calibrated to produce a high-confidence upper bound on the residual, yielding a conservative reachability value used for safety guidance. The robot follows nominal MPC command while shifting toward a reachability-gradient-based safe action near the boundary of the backward reachable tube. Experiments in CrowdSim across in-distribution densities and out-of-distribution marching and circular flows show that proposed method improves navigation efficiency while maintaining competitive emprical safety. Particularly under structured OOD flows, it improves both success and collision rate relative to conformal RL baseline, demonstrating a favorable safety-efficient trade-off.