Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mism...
Yulong Liu, Xiao-Tian Han, Jun-Yuan Shang et al.· 0 citations
Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurrence coverage, or must discriminate competing hypotheses---which can cause failure before substantive reasoning begins. Prescribing a fine-grained solution procedure for...
Can-Can Zhang, Bao-Feng Zhang, Xiao-Tian Han et al.· 0 citations
Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure, and adapts this construction along task and spatial axes to preserve visual information dispersed across frames under compress...
Wen-Ti Yin, Xiao-Tian Han, Jun-Yuan Shang et al.· 0 citations
Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide which tokens or heads to preserve from runtime attention scores, observation windows, calibration prompts, or learned gates, making head diagnos...
YeHan Yang, Jun-Yuan Shang, Yang Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.