Skip to content

Category

artificial intelligence

4,637 papers

QUBRIC: Co-Designing Queries and Rubrics for RL Beyond Verifiable Rewards

QUBRIC, a framework that co-designs queries and rubrics can make rubric-based RL a practical complement to RLVR beyond strictly verifiable tasks, provides evidence that co-designing queries and rubrics can make rubrics a practical complement to RLVR beyond strictly verifiable tasks.

Rongzhi Zhang, Rui Feng, Zhi-Han Zhang et al. · 0 citations

Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

Agentic Chain-of-Thought Steering (ACTS), which formulates reasoning steering as a Markov decision process where a controller agent adaptively steers a frozen reasoner during inference, and enables budget-aware strategy control for efficient reasoning while preserving the reasoner's generation continuity.

Yu Xia, Zhouhang Xie, Xin Xu et al. · 0 citations
#artificial intelligence Review Jun 2026

The Unsampled Truth: Quantifying Prompt Artifacts in LM Psychometrics

When prompting language models for psychometric assessment, researchers assume that the responses reflect the injected persona and the meaning of the survey item. We test this premise using a diagnostic design that crosses five semantically distinct baseline personas with five semantically equivalent variants of each of four prompt components (persona wording, task instruction, item wording, option symbol). Measuring the 1-Wasserstein distance between the resulting response distributions and partitioning the variation among the five components allows for the separation of target effects from prompt artifacts. We apply the framework to 13 open-weight small language models (0.6B to 14B) on the Big Five Inventory and the Short Dark Triad. We find that in most models, the task instruction and option symbol displace response distributions further than paraphrasing the persona description or the item itself. For a substantial share of items, the artifact share of explained variation exceeds 50%; non-semantic changes of the prompt account for more response variation than the baseline personas. Our framework lets researchers quantify these prompt artifacts before interpreting psychometric output.

Nils Schwager, Christoph Hau, Simon Münker et al. · 0 citations

Linguistics-Aware Non-Distortionary LLM Watermarking

LUNA is a linguistically adaptive watermark that combines model-free detection with single-token non-distortion under the standard random-key model, and is the only method that simultaneously achieves AUROC>0.99 and an absolute median perplexity shift below 0.1.

Shinwoo Park, Hyejin Park, Hyeseon An et al. · 0 citations

Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning

SelSkill is proposed, a dual-granularity preference-learning framework for selective skill invocation that formulates skill use as a skill-or-skip decision, uses predictive uncertainty to prioritize candidate decision points, and constructs controlled invoke-skip preference pairs from shared trajectory prefixes.

Chishui Chen, Jiaye Lin, Te Sun et al. · 1 citation · ⚡1

FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding rhetorical hate mechanisms with target community features and preventing causal evaluation of model vulnerabilities. To address this, we introduce FBHM, a systematically curated benchmark of Functionality Based Hateful Memes constructed along two orthogonal axes: 25 distinct rhetorical functionalities and 10 target communities (5,000 memes total). Benchmarking state-of-the-art VLMs reveals a severe generalization gap: models highly accurate on standard datasets catastrophically drop to near-random performance on FBHM, proving they exploit dataset-specific heuristics rather than robust multimodal reasoning. To efficiently close this gap, we propose LSV (learnable steering vectors), an ultra-low data regime strategy that applies a causal intervention objective on as few as 500 steering samples (50 unique base memes), boosting FBHM performance by ~30 Macro-F1 points while outperforming in-context learning and PEFT without degrading source-domain performance.

Paramananda Bhaskar, Naquee Rizwan, Daksh Jogchand et al. · 0 citations

Skill-Conditioned Gated Self-Distillation for LLM Reasoning

This work proposes Skill-Conditioned Gated Gated Self-Distillation (SGSD), which formulates skill-based SD as teacher hypothesis validation rather than unconditional imitation, and shows that SGSD consistently improves over GRPO and remains competitive with answer-conditioned OPSD under a weaker PI assumption.

Jiazhe Huang, Xiao Chen, Xiao Luo et al. · 5 citations

Reverse Probing: Supervised Token-level Uncertainty Quantification for Large Language Models in Clinical Text

Reverse Probing is proposed, the first UQ framework specialized for clinical summarization, which estimates token-level uncertainty directly from pre-existing labeled summaries, and reveals that delta energy and neighborhood context are the most consistent predictors across all models.

Bushi Xiao, Sarvesh Soni, Daisy Zhe Wang · 0 citations
#artificial intelligence Review May 2026

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in German law, is introduced and 12 contemporary LLM systems are evaluated with a rubric-aligned LLM-as-a-Judge cross-validated against a multi-rater human-grading layer.

Sebastian Nagl, A. Mayrhofer, Martin Heidebach et al. · 0 citations

Integrated and Cross-Architecture Interpretation of LLM Reasoning

This work proposes to use bandwidth-calibrated MIP coupled with Tukey IQR peak-detection to isolate reasoning-crucial tokens at the output layer, and applies a Jaccard stability metric over multi-domain problems to verify if the MIP-identified tokens are reasoning quality-guaranteed.

Leonardo Matthew Yauw, Wei-Bin Kou, Yujiu Yang · 0 citations

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

It is found that English-Korean performance gaps vary substantially across models and task families, and that SpokenQA and audio understanding rankings diverge, revealing complementary weaknesses invisible to English-only evaluation.

Haechan Kim, Seung-Jun Chung, Inkyu Park et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.