This work identifies a progressive decay of robust features across network layers and establishes a functional dependency between the prevalence of these features and model performance, and proposes Suppress and Diversify (S\&D), a non-intrusive refinement approach that enhances robustness by dynamically selecting robu...
Jian-Gang Yang, Wen-ku Shi, Xiao-Ran Xu et al.· 0 citations
The model achieves the best performance among similarly sized models on 19 of the 38 benchmarks and substantially outperforms strong competitors, including Qwen3.6-A3B and Cosmos 3.5 MoT-2B, and demonstrates strong performance on embodied agentic tasks requiring multi-turn interaction and long-horizon reasoning.
VTM-Nav re-localizes the agent in accumulated scene structure, retrieves target-relevant records from plausible rooms, and grounds memory guidance in candidates derived from the current observation, demonstrating effective reuse of cross-episode scene experience through hierarchical visual-topological memory.