Large language models (LLMs) have shown strong potential as training-free text encoders for long-context embeddings. Existing approaches primarily improve information flow under causal attention and typically construct embeddings by uniformly averaging all token representations. However, for long documents, such mean p...
Zi-Feng Cheng, Jie Zheng, Zhiwei Jiang et al.· 0 citations
A Dynamic Distribution-Aware Uncertainty Quantification framework (DDA-UQ) is proposed that shifts the paradigm from static mapping to a dynamic distribution-aware process and significantly outperforms state-of-the-art methods.
Ao Zhou, Zhiwei Jiang, Zi-Feng Cheng et al.· 0 citations
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conventional off-policy distillation with substantially less data. However, standard token-level OPD can provide only fragmen...
Changhui Sun, Lan-Bo Liu, Hang Lei et al.· 0 citations
This work proposes a structured suffix modeling method that incorporates the decoding results from the previous step into the suffix token representations at the current step, allowing them to carry evolving denoising information across generation steps.
Zifeng Cheng, Keda Li, Zhiwei Jiang et al.· 0 citations
Class-wise Covariance Regularization is proposed, which aligns the predicted covariance structure of class confidences with the semantic correlations encoded in pretrained text embed-dings with the geometric consistency of the class space throughout fine-tuning, resulting in more stable and interpretable confidence dis...
Ao Zhou, Zhi-Wei Jiang, Zi-Feng Cheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.