Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field...
Shao-Jun Xia, Hui-Xin Zhang, Zhen Lei et al.· 0 citations
ARSTAG is presented, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data and shows that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size.
This work proposes ChronoVision, a multimodal framework designed to align visual logic with latent imagery, and introduces Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task.
This work introduces Model-Based Diffusion via Constraint Optimization and Adaptive Scheduling (MD-COAS) for SRMP that unifies the inexact Augmented Lagrangian Method (iALM) soft diffusion prior with a Convex Feasible Set (CFS)-based hard projection operator, and adaptively schedules and co-optimizes safety enforcement...