Semantic Flow Regularization (SFR), a lightweight auxiliary objective that supervises the backbone with continuous sentence-encoder embeddings of future segments via conditional flow matching, improves output diversity, style fidelity, and response quality over SFT on a large-scale industrial dialogue dataset.
Ke Peng, Feifei Li, Xing Fan et al.· arXiv.org· 0 citations
Evaluating three frontier VLMs in both homogeneous and cross-model adversarial settings, it is found that even the strongest agent hallucinates 15.1% of its verifiable spatial claims and 11.5% of accusations are strictly unsupported.
Ye Yuan, Ruiqi Song, Wei-En Li et al.· arXiv.org· 2 citations
DRIP-R is introduced, a benchmark that systematically exploits real-world retail policy ambiguities to construct scenarios in which no single correct resolution exists, and shows that frontier models fundamentally disagree on identical policy-ambiguous scenarios, confirming that ambiguity poses a genuine and systematic challenge to LLM decision-making.
Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang et al.· arXiv.org· 0 citations
This paper introduces KoALa-Bench, a comprehensive benchmark for evaluating Korean speech understanding and speech faithfulness of LALMs, and incorporates listening questions from the Korean college scholastic ability test as well as content covering Korean cultural domains.
Jinyoung Kim, Hyeongsoo Lim, Eunseo Seo et al.· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
LLM annotation at scale outperforms human-supervised classifiers at roughly one-tenth the cost, for both a closed-source and an open-weight LLM, and the advantage is robust under soft-label evaluation.
Ahmad Dawar Hakimi, Lea Hirlimann, Isabelle Augenstein et al.· 0 citations
A systematic analysis of expert routing patterns in MoE models reveals Language Routing Isolation, in which high- and low-resource languages tend to activate largely disjoint expert sets, and proposes RISE, a framework that exploits routing isolation to identify and adapt language-specific expert subnetworks.
Kening Zheng, Wei-Chieh Huang, Jiahao Huo et al.· arXiv.org· 4 citations· ⚡2
NeuRIT is proposed, a Neuron-guided Robust Instruction-Tuning framework built on a localization-first perspective that mines context-aware neurons associated with relevant and irrelevant context processing, and uses them as anchors to selectively adapt both the identified neuron groups and the layers in which they concentrate.
Jae Lee, Jaemin Kim, Sumyeong Ahn et al.· 0 citations
This study tested five Large Language Models and compared their performance to that of human controls using an adapted version of a text-based tool widely used in human ToM research, revealing a performance gap between the models.
Anna Babarczy, András Lukács, Péter Vedres et al.· arXiv.org· 1 citation
Experimental results show that PEFT consistently strengthens hallucination detection ability, substantially improving AUROC across a wide range of hallucination detectors, and indicates that PEFT methods primarily reshapes how uncertainty is encoded and surfaced, comparing with injecting new factual knowledge into the models.
It is taken that TLMs encode a non-trivial amount of syntactic knowledge, which shows strong performance on formal syntactic phenomena, but weaker and more variable performance on phenomena at the syntax-semantics interface.
This work introduces MentorQA, the first multilingual dataset and evaluation framework for mentorship-focused question answering from long-form videos, and defines mentorship-focused evaluation dimensions that go beyond factual accuracy, capturing clarity, alignment, and learning value.
Parth Bhalerao, D. D’souza, Rui Guan et al.· arXiv.org· 0 citations