Long-horizon camera-controlled video generation requires recovering previously observed content from an ever-growing visual history. Existing approaches either search historical context implicitly or reconstruct it into persistent 3D memory, facing inefficient memory access or accumulated geometric errors. Our key insi...
Ze-Song Yang, Wei-Kai Chen, Li-Yuan Cui et al.· 0 citations
CoverPrune is introduced, a training-free framework that formulates inference-time token pruning as an Optimal Transport (OT) problem, and the Feature-Spatial-Temporal (FST) transport cost and target capacity, along with an efficient Spatial-Guided Greedy Selection (SGS) algorithm to approximate the OT objective.
Peng Ling, Ying-Da Yin, Ling-Ting Zhu et al.· 0 citations
A framework that predicts an executable scene representation from sparse views and realizes it through perception-grounded asset instantiation and closed-loop spatial refinement is presented, and observation-consistent supervision is introduced that aligns each target scene with the visual evidence available in its inp...
Kai Li, Lu-Tao Jiang, Zhenyang Li et al.· 1 citation
The key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling and replacing global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the...