EcoVLA: Energy-Efficient Device-Edge Co-inference for Vision-Language-Action Models Under Real-Time Constraints
Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inference cost poses significant challenges for deployment in robotic systems. In practice, on-device inference is constrained by limited compute capacity and energy budgets, struggling to simultaneously satisfy r...