Skip to content

Author

Yi-Fan Wang

We have 6 of 27 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also anticipating where the interaction is heading. Existing policies largely leave such foresight to emerge implicitly from act...

Zhong-Bo Zhang, Zai-Bin Zhang, Yi-Fan Wang et al. · 0 citations
Preprint Oct 2026

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models

Generalization in multi-arm collaboration can be studied as composing familiar atomic skills in new ways across arms. However, existing evaluations offer limited insight into which training and architectural choices support this ability under different coordination requirements. We introduce \textbf{ACG-Bench}, a bench...

Zai-Bin Zhang, Bing-Hao Ran, Yu-Han Wu et al. · 0 citations
Preprint Sep 2026

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may be partially occluded or tightly intermingled with visually similar distractors. As a result, standard 3D diffusion policies often struggle to localize and exploit task-relevant geometry as scene com...

Chang-Bo Yan, Zhong-Bo Zhang, Zai-Bin Zhang et al. · 0 citations
Preprint Sep 2026

MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs

MinCU is introduced, a benchmark for grounded minimal-change understanding and Semantic-Guided Implicit Spatial Anchors (SG-ISA), a structured autoregressive method that decomposes prediction into a Think-Locate-Describe sequence, suggesting that an implicit intermediate spatial interface can be more effective than rel...

Chao-Qian Mu, Wen-Hao Wu, Zi-Chen Liang et al. · 0 citations
Preprint Sep 2026

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

This work proposes CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution, and introduces the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anoma...

Jun-Lan Xiao, Jun-Wei Jiang, Zai-Bin Zhang et al. · 0 citations
Preprint Aug 2026

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

MA-VLA decomposes cooperative behavior into mid-level atomic prompts and allocates them to individual arms, enabling explicit subgoal specification and compositional reuse across tasks and indicates that structured, per-arm atomic action assignment offers a practical route to scalable generalization in multi-arm embodi...

Zai-Bin Zhang, Jun-Lan Xiao, Zhong-Bo Zhang et al. · 4 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.