Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage a...
Wen-Xiao Fan, Jing-Ling Fu, Li-Chen Ma et al.· 0 citations
TransAnyText, a structured visual code framework that reformulates image text translation as generating renderable HTML patches from source images and target languages, is introduced, providing an effective, controllable, and editable solution for cross-border e-commerce image translation.
Xiao-An Liu, Li-Chen Ma, Zipeng Guo et al.· 0 citations
Observation Operator Diffusion is proposed, a unified framework that aligns both the supervision trajectory and feature refinement with the intrinsic recovery order of image structures and introduces GL-CoDA, a decoder that injects scale-specific Gaussian-Lanczos observations across decoding stages for coarse-to-fine f...
Shaojie Guo, Li-Chen Ma, Haoyang Tong et al.· 1 citation
Energy-Guided Flow Matching is introduced that explicitly models a coarse-to-fine generative trajectory by moving endpoint that evolves smoothly from low-frequency image to clean image and requires no adaptation of the backbone and training data.
Haoyang Tong, Yu He, Fang Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.