Jul 2026
FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation
FM-VLA is proposed, a VLA model with force-based memory, enabling temporal context reasoning for non-Markovian, contact-rich manipulation, and achieves over 80% success rate with minimal inference overhead, significantly outperforming baseline approaches.
Ruicheng Li, Qi-Xiu Li, Ruichun Ma et al.
· arXiv.org · 1 citation