Skip to content

Author

Vincent-Daniel Yun

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving

KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history. Under aggressive budgets, this not only lowers accuracy but can also cause runaway degeneration, where the model produces incoherent or repetitive tokens until reaching t...

Minsoo Cheong, Woo-Sang Lim, Vincent-Daniel Yun et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Output-Aware Rotation for INT2 KV-Cache Quantization

The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. However, existing rotation-based INT2 methods optimize cache statistics or proxy errors before the complete attention readout, even though...

Vincent-Daniel Yun, Woo-Sang Lim, Minsoo Cheong et al. · 1 citation

Locality-Aware Redundancy Pruning for LLM Depth Compression

It is shown that inter-layer redundancy can be either localized or globally distributed depending on the LLM architecture, and Representation Locality Score (RLS) is introduced, derived from global inter-layer hidden-state similarity.

Vincent-Daniel Yun, Youngrae Kim, Woosang Lim et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.