Jun 2026· arXiv.org· Vol abs/2606.05616· 0 citations· 32 references
Computer Science
TL;DR
This work introduces a framework for identifying whether a model's drug semantics are driven mainly by the affix, the stem, or the drug name as a whole, and reveals that models often induce drug meaning primarily through affix cues, yet rarely explicitly indicate this reliance, and sometimes incorrectly conflate properties among affix-sharing drugs.
Abstract
The morphological form of a word can often give cues to its meaning, but purely relying on these mappings can lead to overgeneralization in high-stakes domains. In the medical domain, for instance, LLMs can confidently reason about fictitious drugs from their affixes alone (e.g., wugcillin) and generate plausible-looking clinical content. We present a behavioral and mechanistic study of LLM"affix heuristics"in pharmacology. Using fictitious drug names built from real affixes, we show that affix signals alone elicit class-level pharmacological responses. We introduce a framework for identifying whether a model's drug semantics are driven mainly by the affix, the stem, or the drug name as a whole. Applied across 653 drugs, our framework reveals that models often induce drug meaning primarily through affix cues, yet rarely explicitly indicate this reliance, and sometimes incorrectly conflate properties among affix-sharing drugs. Activation patching across models further localizes this behavior to early-mid layers. These findings show that morphological shortcuts pose a subtle but measurable risk to safety.
The capacity to retrieve semantic information, particularly related to unique entities, is thought to depend on the anterior temporal lobe (ATL). Evidence underpinning the ATL's characterization as a hub of the semantic network has depended largely on naming tasks requiring lexical retrieval. Consequently, questions ab...
Deom Nicolas, Khalil Omar, Protzner Andrea et al.· Neuropsychologia· 0 citations
CNeo-Bench, a benchmark of 4,759 Chinese neologisms with reference definitions, is introduced, organized into five top-level categories and nine subcategories by the linguistic mechanism behind each expression, paired with a two-tier evaluation framework that separates whether a model can describe a neologism from whet...
Kai-Yan Zhao, Zhong-Tao Miao, Zhe-Yong Xie et al.· 0 citations
Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or"probes") are widely used for this task, they face significant methodological criticism: training auxiliary classifiers introduces capacity confounds...
Language is vastly ambiguous, yet in context our brains resolve ambiguity effortlessly. This is particularly evident at the lexical level: Although most words can convey a wide range of related meanings, humans rapidly and automatically select the one appropriate for the context. We investigated the neural mechanisms u...
Simone M. Krogh, L. Pylkkänen· bioRxiv· 0 citations
The Malay suffix ‑
an
plays a central role in the language’s system of derivational morphology, and we argue that its role extends far beyond simple nominalization. This paper examines the semantic range of ‑
an
and its metaphorical and relational networks based on previous linguistic research and semantic anal...
Siaw-Fong Chung· International Review of Prag...· 0 citations
It is found that while all models show sensitivity to existential presupposition across syntactic embeddings, determiner types and contextual cues, their behaviour differs markedly in strength and systematicity, with NLI-fine-tuned autoregressive models exhibiting the most coherent and stable projection patterns.
Marie-Léontine Wörgötter, Shiyang Lai, Sebastian Schuster· International Conference on...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.