Egocentric cameras are widely used in robotic navigation and manipulation, yet conventional 2D Video Object Tracking (VOT) methods suffer from severe performance degradation under rapid viewpoint changes and frequent frame-out events. Because most existing trackers rely solely on 2D appearance cues, they often fail to...
Hayeoung You, Sangbeom Lee, Huisu Kim et al.· 2026 23rd International Conf...· 0 citations
How do humans navigate to a target object in an unmapped, unseen environment? We certainly do not wander aimlessly. Instead, human explorers naturally rely on spatial context, leveraging the inherent co-occurrence of everyday objects to infer a target’s probable location. However, conventional zero-shot object navigati...
Sangmin Park, Minhwan Ko, Kyoobin Lee· 2026 23rd International Conf...· 0 citations
Vision-Language-Action (VLA) models have demonstrated strong performance in robot manipulation by leveraging pre-trained vision-language models to map observations directly to actions. However, existing approaches reason primarily at the visual or semantic level, lacking explicit understanding of the physical interacti...
Kangmin Kim, Geonhyup Lee, Sangbeom Lee et al.· 2026 23rd International Conf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.