CoViST: Visual Token Compression via Composable States
CoViST is proposed, a training-free framework that represents a compressed image as a composable visual state that retains its interpretation under successive reductions, enabling the same formulation to support both fixed compression before prefill and progressive compression within the decoder.