This work proposes using Tensor Product Representations (TPRs) as a unifying hypothesis, and shows that TPRs can unify several prior interpretability methods: additive analogies, linear probing, sparse autoencoders, and activation patching.
Evaluating context mechanisms requires sequence-aligned term-level, insertion, and speaker-label measures alongside aggregate accuracy, and sequence-alignment analysis found a small improvement on complete context-listed phrases, too small to materially change side-level WER, and for Gemini coexisting with worsened unlisted-token error.
Theodore O. Cochran, Stephanie Dodson, Keith Nore· 0 citations
GreenBench, a benchmarking framework that evaluates the energy efficiency, throughput, and carbon footprint of five open-source LLMs across three NLP tasks on an Apple M4 Pro with 48 GB unified memory, is presented.
R. Kannan, Rajendra P. Firke, Shreya Bengle et al.· 0 citations
This evaluation shows that while LLMs effectively uncover a substantial portion of implicitly-related touchpoints, significant room for improvement remains in their selection performance, and offers a new roadmap for transitioning conversion attribution from mechanical rule-matching to human-aligned semantic reasoning.
Jinqi Wu, Sishuo Chen, Zhangming Chan et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Within low resource domains, results identify the model's parsing of information and subsequent reasoning as the source of reasoning failure, rather than corpus contents, rather than corpus contents.
CLAIMPROBE is introduced, a claim-level audit that decomposes DR reports into claims and measures hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence and proposes CLAIMWRITER, a hierarchical claim-based writer that extracts source facts, maps them to a query-derived outline, and drafts each section from a source-linked claim representation.
Hiroaki Hayashi, P. Venkit, Prafulla Kumar Choubey et al.· 0 citations
Evaluation of six frontier models reveals that even the strongest model reaches only 63.1\% pass rate, with many tasks unsolved by any model, highlighting that multilingual coding competence is a distinct and underexplored capability axis.
Yunsu Kim, Kaden Uhlig, Ashwin Purohit et al.· 0 citations
The Prompt Phrases Prediction Network (PPN) is introduced, an encoder-decoder architecture designed to effectively extract keyword prompts embeddings and infuse the prompt embedding into the Prompt-guided KWS encoder by utilizing a Prompt-acoustic Multi-head Cross-attention (MHCA).
G. Xu, Cheng-Fei Li, Xian-Liang Wang et al.· 0 citations
The results indicate that current MLLMs can process visually clear English and structured numeric content, but reliable native Khmer document understanding remains an open challenge.
PAUSE (Pause-And-Update Strategy Editing) is an intervention that exposes an editable adaptation strategy as a human control surface for cultural decisions in long-form story adaptation, a structured artifact that can be inspected, edited, and then projected through downstream character, entity, and chapter-localization stages.
Taaha Kazi, Vasu Sharma, M. Saifullah et al.· 0 citations
This work proposes LPS-TC, a Lightweight Proactive Speech Turn Controller for plug-and-play integration, and introduces a two-tier evaluation scheme that assesses both chunk-level timing precision and turn-level interaction quality under realistic streaming constraints.
Tianrui Pan, Qinglin Zhang, Chong Deng et al.· 0 citations
This study establishes an end-to-end prototype from raw BIM data input, through defect identification, to repair suggestion generation, and establishes an end-to-end prototype to identify and repair various defects in BIM via domain-specific LLMs.
Jia-Rui Lin, Yunzhen Cai, Xiang Ni et al.· 0 citations