QUBRIC, a framework that co-designs queries and rubrics can make rubric-based RL a practical complement to RLVR beyond strictly verifiable tasks, provides evidence that co-designing queries and rubrics can make rubrics a practical complement to RLVR beyond strictly verifiable tasks.
Agentic Chain-of-Thought Steering (ACTS), which formulates reasoning steering as a Markov decision process where a controller agent adaptively steers a frozen reasoner during inference, and enables budget-aware strategy control for efficient reasoning while preserving the reasoner's generation continuity.
When prompting language models for psychometric assessment, researchers assume that the responses reflect the injected persona and the meaning of the survey item. We test this premise using a diagnostic design that crosses five semantically distinct baseline personas with five semantically equivalent variants of each of four prompt components (persona wording, task instruction, item wording, option symbol). Measuring the 1-Wasserstein distance between the resulting response distributions and partitioning the variation among the five components allows for the separation of target effects from prompt artifacts. We apply the framework to 13 open-weight small language models (0.6B to 14B) on the Big Five Inventory and the Short Dark Triad. We find that in most models, the task instruction and option symbol displace response distributions further than paraphrasing the persona description or the item itself. For a substantial share of items, the artifact share of explained variation exceeds 50%; non-semantic changes of the prompt account for more response variation than the baseline personas. Our framework lets researchers quantify these prompt artifacts before interpreting psychometric output.
Nils Schwager, Christoph Hau, Simon Münker et al.· 0 citations
LUNA is a linguistically adaptive watermark that combines model-free detection with single-token non-distortion under the standard random-key model, and is the only method that simultaneously achieves AUROC>0.99 and an absolute median perplexity shift below 0.1.
Shinwoo Park, Hyejin Park, Hyeseon An et al.· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
SelSkill is proposed, a dual-granularity preference-learning framework for selective skill invocation that formulates skill use as a skill-or-skip decision, uses predictive uncertainty to prioritize candidate decision points, and constructs controlled invoke-skip preference pairs from shared trajectory prefixes.
Chishui Chen, Jiaye Lin, Te Sun et al.· arXiv.org· 1 citation· ⚡1
Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding rhetorical hate mechanisms with target community features and preventing causal evaluation of model vulnerabilities. To address this, we introduce FBHM, a systematically curated benchmark of Functionality Based Hateful Memes constructed along two orthogonal axes: 25 distinct rhetorical functionalities and 10 target communities (5,000 memes total). Benchmarking state-of-the-art VLMs reveals a severe generalization gap: models highly accurate on standard datasets catastrophically drop to near-random performance on FBHM, proving they exploit dataset-specific heuristics rather than robust multimodal reasoning. To efficiently close this gap, we propose LSV (learnable steering vectors), an ultra-low data regime strategy that applies a causal intervention objective on as few as 500 steering samples (50 unique base memes), boosting FBHM performance by ~30 Macro-F1 points while outperforming in-context learning and PEFT without degrading source-domain performance.
This work introduces AgentREVEAL, a diagnostic framework for analyzing retrieval-induced safety degradation in LLM agents, and uncovers the Safe Source Paradox, a safety-utility trade-off for retrieval-enabled agents.
Aditya Nawal, Manit Baser, M. Gurusamy· arXiv.org· 0 citations
This work proposes Skill-Conditioned Gated Gated Self-Distillation (SGSD), which formulates skill-based SD as teacher hypothesis validation rather than unconditional imitation, and shows that SGSD consistently improves over GRPO and remains competitive with answer-conditioned OPSD under a weaker PI assumption.
Jiazhe Huang, Xiao Chen, Xiao Luo et al.· arXiv.org· 5 citations
Reverse Probing is proposed, the first UQ framework specialized for clinical summarization, which estimates token-level uncertainty directly from pre-existing labeled summaries, and reveals that delta energy and neighborhood context are the most consistent predictors across all models.
BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in German law, is introduced and 12 contemporary LLM systems are evaluated with a rubric-aligned LLM-as-a-Judge cross-validated against a multi-rater human-grading layer.
Sebastian Nagl, A. Mayrhofer, Martin Heidebach et al.· arXiv.org· 0 citations
This work proposes to use bandwidth-calibrated MIP coupled with Tukey IQR peak-detection to isolate reasoning-crucial tokens at the output layer, and applies a Jaccard stability metric over multi-domain problems to verify if the MIP-identified tokens are reasoning quality-guaranteed.
Leonardo Matthew Yauw, Wei-Bin Kou, Yujiu Yang· arXiv.org· 0 citations
It is found that English-Korean performance gaps vary substantially across models and task families, and that SpokenQA and audio understanding rankings diverge, revealing complementary weaknesses invisible to English-only evaluation.
Haechan Kim, Seung-Jun Chung, Inkyu Park et al.· arXiv.org· 0 citations