Skip to content

Author

Junyuan Shang

We have 6 of 20 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mism...

Yulong Liu, Xiao-Tian Han, Jun-Yuan Shang et al. · 0 citations
Preprint Aug 2026

From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents

Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurrence coverage, or must discriminate competing hypotheses---which can cause failure before substantive reasoning begins. Prescribing a fine-grained solution procedure for...

Can-Can Zhang, Bao-Feng Zhang, Xiao-Tian Han et al. · 0 citations
Preprint Aug 2026

Aggregating Visual Information with Optimal Transport for VideoLM Token Compression

Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure, and adapts this construction along task and spatial axes to preserve visual information dispersed across frames under compress...

Wen-Ti Yin, Xiao-Tian Han, Jun-Yuan Shang et al. · 0 citations
Preprint Aug 2026

Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry

Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide which tokens or heads to preserve from runtime attention scores, observation windows, calibration prompts, or learned gates, making head diagnos...

YeHan Yang, Jun-Yuan Shang, Yang Li et al. · 0 citations

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

This work presents a principled analysis of the distributions induced by lossy verification methods, and shows that many seemingly distinct approaches differ only superficially and can be unified into two categories: truncation-based verification and collaborative verification.

Tian-Yu Wang, Yuxuan Zhou, Heng Li et al. · 0 citations
Preprint Aug 2026

Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

Memory-Augmented Compression is proposed, a training-free framework that constructs reusable reasoning memories from historical traces and retrieves them as prefill-side scaffolds to compensate for information lost during compression.

Si-Meng Zhang, Yi-Long Chen, Wenyuan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.