Skip to content

Category

artificial intelligence

6,270 papers

#artificial intelligence Preprint Aug 2026

Writing Style Similarity Reflects Academic Genealogy

A corpus of arXiv authors with solo papers from the Mathematics Genealogy Project graph is built, giving 5 total authors and ground-truth advisor-student pairings, where advisors sit closer in cosine distance to their students than a random same-field author does.

Cameron Manzo · 0 citations
#artificial intelligence Preprint Aug 2026

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

Comprehensive in-domain and out-of-domain evaluations across multiple LLMs demonstrate that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines and reveal improvements in self-correction frequency and effectiveness.

Duc Anh Vu, N. Hoang, Do Xuan Long et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse

Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks mainly annotate surface polarity or final emotion categories, while lacking a structured account of how explicit expression, implicit affect, pragmatic intent, and fine grained emotion interact. This limitation makes current evaluations insensitive to cases where affective meaning is concealed, weakened, inverted, or pragmatically reshaped, thereby obscuring model failures in deeper emotion understanding. To address this gap, we introduce CUE Bench, a Chinese Unsaid Emotion benchmark that centers on Affective Stance and covers diverse communicative scenarios. CUE Bench constructs nine human interpretable affective stances from explicit implicit polarity interaction and further provides intent and fine grained emotion annotations for structured affective inference. Experiments show that incorporating Affective Stance improves fine grained emotion recognition by 3.1 percentage points and pragmatic intent detection by 8.1 percentage points over strong baselines.

Zhen-Yan Zheng, Yun-Yao Zhang, Jun Sheng et al. · 0 citations
#artificial intelligence Review Aug 2026

Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics

This work proposes Sampled-BPE, a lightweight token-level auditing pipeline that sample a small subset and train BPE tokenizer to surface polluted tokens, and releases a hierarchical Chinese web token dataset with 660k+ token records, organized as trees to support review and tracing of pollution.

Qingjie Zhang, Ziqi Tang, Jie Zhang et al. · 0 citations
#artificial intelligence Review Aug 2026

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

Evaluating six matched base and post-trained models on the Pew American Trends Panel, it is found that base models are the stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure.

Seth Grief-Albert, Jessica Y. Bo, Di-Fan Jiao et al. · 0 citations
#artificial intelligence Review Jul 2026

SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation

Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of annotations, rendering human labeling prohibitively costly. While recent work has demonstrated synthetic label generation using LLMs, deploying such approaches at industrial scale requires integrated quality control mechanisms. We present SynthAVE, a large-scale human-validated benchmark for attribute value extraction spanning 12,726 products across 229 product types, 792 attributes, and 4 languages (Spanish, French, Italian, German). To validate synthetic labels at scale, we introduce a multi-LLM arena framework where samples are independently evaluated by 21 judge configurations (7 model families $\times$ 3 prompts), with final labels determined via majority voting. The majority vote ensemble agrees with human experts at Cohen's $\kappa = 0.92$ (95.2% agreement), while individual judges show substantial inter-model agreement (Fleiss'$\kappa = 0.76$). This demonstrates that diverse models with varying individual judgments aggregate into highly reliable predictions, enabling cost-effective validation at scale while maintaining quality parity with human review.

Andrea Scarinci, V. Negri, Brayan Impata et al. · 0 citations
#artificial intelligence Review Jun 2026

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

PhysAssistBench is introduced, a benchmark for interactive doctor-patient-EHR assistance that uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality.

T. Du, Peijie Yu, Sihan Shang et al. · 0 citations

Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

ADAS is proposed, a training-free reranking rule that leaves the base sampler's stopping rule unchanged and greedily discounts each token-wise confidence score according to its attention to already selected positions, weighted by their prediction uncertainty.

Y. Şahin, Ahmed R. Saikia, V. Cevher et al. · 1 citation

Summarization is Not Dead Yet

A multi-track evaluation covering diverse datasets and state-of-the-art LLMs reveals a more nuanced landscape in which human references continue to demonstrate advantages in informativeness and faithfulness, whereas LLM outputs are preferred mainly for surface-level coherence and fluency.

Dongqi Liu, Chenxi Whitehouse, Zheng Zhao et al. · 0 citations

ABLE: Representing and Mapping LLMs via Attribution-Based Large-model Embedding

ABLE (Attribution-Based Large-model Embedding), a framework that leverages the interpretability space to construct model representations by aggregating gradient-based feature attributions via a tokenizer-agnostic word-level alignment, captures model-specific input-sensitivity patterns rather than only surface-level outputs.

Zirui Wang, Yusen Hou, Shaofeng Liang et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.