Reasoning-Distilled Planning: Transferring Vision–Language Guidance to Real-Time End-to-End Driving
Abstract
Real-time end-to-end planners learn metric trajectories efficiently, but logged waypoints provide little supervision about scene intent in ambiguous interactions. We introduce Reasoning-Distilled Planner (RDP), which uses a vision–language model only while constructing training signals for a sparse planner. The teacher returns schema-constrained scene reasoning and an auxiliary trajectory. Deterministic parsing, repair, geometry tests, and agreement checks decide which parts of that output can supervise the student. Expert waypoints continue to anchor optimization; valid teacher trajectories receive confidence- and consistency-dependent weight, and teacher tokens are transferred through a marginal Wasserstein alignment that requires no token-to-query pairing. Deployment discards the complete teacher path. On nuScenes, the resulting planner runs at 7.0 FPS with 0.54 m mean L2 waypoint error and a 0.04% mean collision rate.