Skip to content
Review

Structuring Semantic Embeddings for Principle Evaluation: A Prototype-Guided Contrastive Learning Approach

Aug 2026 · 1 citation · 54 references
Computer Science

TL;DR

This paper introduces Prototype-Guided Contrastive Learning (PGCL), a prototype-guided geometric regularization module built on top of frozen text embeddings, and introduces Prototype-Guided Contrastive Learning (PGCL), a prototype-guided geometric regularization module built on top of frozen text embeddings.

Abstract

Reliable post-hoc evaluation asks whether already generated text satisfies a target criterion after generation. In this paper we study a focused frozen-embedding setting using principle-evaluation proxy tasks: toxicity detection, fine-grained emotion categorization, and ordinal review rating. General-purpose text embeddings are widely deployed for such tasks, but broad semantic similarity can place semantically similar yet task-distinct examples in overlapping regions of the representation space. We introduce Prototype-Guided Contrastive Learning (PGCL), a prototype-guided geometric regularization module built on top of frozen text embeddings. The module combines a semantic stream, a prototype-anchor attention stream, supervised contrastive learning, offset-based prototype-margin regularization, and stream regularization to produce a compact task-adapted representation without updating the base encoder. Controlled experiments show that PGCL improves over raw frozen embeddings on all three datasets and gives the clearest direct-baseline margin on AmazonReviews, while remaining competitive with strong direct frozen metric-learning baselines on GoEmotions and ToxicComment. We also add supervised residual-adapter, encoder-LoRA, full fine-tuning, objective ablation, sensitivity, and fully logged few-shot LLM protocol diagnostics to define the boundary of the claim. The theoretical analysis is revised as a sufficient-condition account for prototype-margin behavior under explicit assumptions in the prototype-mapping space, rather than as an unconditional training or final-embedding separation guarantee.

View source

Similar papers

Book Open access Feb 2026

Distribution-Level Contrastive Supervision for Generative Recommendation

SODA, a plug-and-play alignment framework that adopts a BPR-style contrastive objective to align recommender representations with target-side distributional representations against negative ones, is developed and demonstrated that SODA consistently strengthens diverse generative recommendation architectures.

Zi-Qiu Xue, Ding-Xian Wang, Yi-Meng Bai et al. · 0 citations
#artificial intelligence Preprint Aug 2026

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

UMER replaces item-wise reflection with Pair-Aware Discriminative Reasoning, which compares query--candidate pairs to identify instruction-relevant matching and discrepancy evidence and achieves state-of-the-art performance under comparable experimental settings while supporting budget-adjustable inference.

Libiao Chen, Xiyang Liu, Yanheng Wei et al. · 1 citation
Open access Aug 2026

A JEPA-Inspired Span-Masked Framework for Language Representation Learning: Revisiting Cosine Similarity and VICReg Regularization

Joint-Embedding Predictive Architectures (JEPAs) have recently emerged as a promising paradigm for self-supervised representation learning by predicting latent embeddings from partial observations rather than reconstructing raw inputs. Although JEPA has demonstrated considerable success in computer vision and large-sca...

C. Jareanpon, Khanabhorn Kawattikul · 0 citations
#artificial intelligence Preprint Aug 2026

Learning from What You Retrieve: Online RL Fine-Tuning for Semantic Retrieval

This work proposes PAO (Positive-Advantage-Only), a selective RL optimization method that selectively applies gradient updates only to retrieved items with positive advantages, effectively pulling query embed- dings toward high-reward regions while preserving global topo- logical stability.

Shao-Wei Wei, Chong Huang, Songtao Fang et al. · 0 citations
Preprint Aug 2026

Douyin Multimodal Embedding Model Technical Report

The Douyin Multimodal Embedding (DME), a model trained in two stages to combine both strengths, is presented, a model trained in two stages to combine both strengths.

Hao-Nan Chen, Chu Li, Zhi-Cheng Wang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.