Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusion-based drafters further reduce drafting latency by proposing multiple tokens at once. However, these tokens are predicted independently, so a single early error causes...
Long-Xuan Yu, Bingsen Chen, Peng Shi et al.· 0 citations
A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively write...
Xi-Rui Li, Peng Shi, Ming-Wen Dong et al.· 0 citations
Vorch-Omni is presented, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation that supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video tr...
Vorch Team, Xiaoyu Chen, Yang Ding et al.· 0 citations
This paper proposes MobileDreamer, an efficient world-model-based lookahead framework to equip the GUI agents based on the future imagination provided by the world model, which consists of textual sketch world model and rollout imagination for GUI agent.
This survey reviews Multimodal Code Intelligence, covering systems that generate, edit, refine, or reason with code under visually grounded inputs and outputs and organizes benchmarks and methods into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Framework...