A corpus of arXiv authors with solo papers from the Mathematics Genealogy Project graph is built, giving 5 total authors and ground-truth advisor-student pairings, where advisors sit closer in cosine distance to their students than a random same-field author does.
Comprehensive in-domain and out-of-domain evaluations across multiple LLMs demonstrate that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines and reveal improvements in self-correction frequency and effectiveness.
Duc Anh Vu, N. Hoang, Do Xuan Long et al.· 1 citation
Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks mainly annotate surface polarity or final emotion categories, while lacking a structured account of how explicit expression, implicit affect, pragmatic intent, and fine grained emotion interact. This limitation makes current evaluations insensitive to cases where affective meaning is concealed, weakened, inverted, or pragmatically reshaped, thereby obscuring model failures in deeper emotion understanding. To address this gap, we introduce CUE Bench, a Chinese Unsaid Emotion benchmark that centers on Affective Stance and covers diverse communicative scenarios. CUE Bench constructs nine human interpretable affective stances from explicit implicit polarity interaction and further provides intent and fine grained emotion annotations for structured affective inference. Experiments show that incorporating Affective Stance improves fine grained emotion recognition by 3.1 percentage points and pragmatic intent detection by 8.1 percentage points over strong baselines.
Zhen-Yan Zheng, Yun-Yao Zhang, Jun Sheng et al.· 0 citations
This work proposes Sampled-BPE, a lightweight token-level auditing pipeline that sample a small subset and train BPE tokenizer to surface polluted tokens, and releases a hierarchical Chinese web token dataset with 660k+ token records, organized as trees to support review and tracing of pollution.
Qingjie Zhang, Ziqi Tang, Jie Zhang et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Evaluating six matched base and post-trained models on the Pew American Trends Panel, it is found that base models are the stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure.
Seth Grief-Albert, Jessica Y. Bo, Di-Fan Jiao et al.· 0 citations
The results show that medical sycophancy depends as much on how a model is challenged and evaluated as on which model is tested, and fabricated evidence has opposite effects across interaction structures.
Kaike Ping, Buse Çarik, Caleb Wohn et al.· 2 citations
Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of annotations, rendering human labeling prohibitively costly. While recent work has demonstrated synthetic label generation using LLMs, deploying such approaches at industrial scale requires integrated quality control mechanisms. We present SynthAVE, a large-scale human-validated benchmark for attribute value extraction spanning 12,726 products across 229 product types, 792 attributes, and 4 languages (Spanish, French, Italian, German). To validate synthetic labels at scale, we introduce a multi-LLM arena framework where samples are independently evaluated by 21 judge configurations (7 model families $\times$ 3 prompts), with final labels determined via majority voting. The majority vote ensemble agrees with human experts at Cohen's $\kappa = 0.92$ (95.2% agreement), while individual judges show substantial inter-model agreement (Fleiss'$\kappa = 0.76$). This demonstrates that diverse models with varying individual judgments aggregate into highly reliable predictions, enabling cost-effective validation at scale while maintaining quality parity with human review.
Andrea Scarinci, V. Negri, Brayan Impata et al.· 0 citations
PhysAssistBench is introduced, a benchmark for interactive doctor-patient-EHR assistance that uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality.
T. Du, Peijie Yu, Sihan Shang et al.· arXiv.org· 0 citations
BioDivergence offers a more faithful way to distinguish contextual divergence from direct contradiction and to separate article-level memorization from genuine task learning.
Elias Hossain, S. S. Jennifer, Sabera Akter Bushra et al.· arXiv.org· 0 citations
ADAS is proposed, a training-free reranking rule that leaves the base sampler's stopping rule unchanged and greedily discounts each token-wise confidence score according to its attention to already selected positions, weighted by their prediction uncertainty.
Y. Şahin, Ahmed R. Saikia, V. Cevher et al.· arXiv.org· 1 citation
A multi-track evaluation covering diverse datasets and state-of-the-art LLMs reveals a more nuanced landscape in which human references continue to demonstrate advantages in informativeness and faithfulness, whereas LLM outputs are preferred mainly for surface-level coherence and fluency.
ABLE (Attribution-Based Large-model Embedding), a framework that leverages the interpretability space to construct model representations by aggregating gradient-based feature attributions via a tokenizer-agnostic word-level alignment, captures model-specific input-sensitivity patterns rather than only surface-level outputs.