2026· International Conference on Language Resources and Evaluation· pp. 9745-9755· 0 citations· 28 references
Computer Science
TL;DR
A layer-wise analysis indicates that surface-level features such as temporality and negation are captured more reliably than deeper semantic phenomena like quantification in large language models, highlighting the limited capacity of current LLMs to generate fully formal meaning representations.
Abstract
We evaluate large language models (LLMs) through semantic parsing into Yarn, a structured meaning representation that distinguishes predicate–argument structure from higher-level linguistic features such as tense, aspect, and modality. For evaluation, we employ SmatchY, a fine-grained metric designed to assess different layers of meaning independently. Our experiments test multiple LLMs under varied conditions, including inference modes, linearization formats (JSON and logic-inspired CFG), and the presence or absence of auxiliary supervision via partial semantic parses. Results show that model performance is highly sensitive to both representational design and supervision, with no single configuration consistently outperforming the others. While some models gain from additional semantic information in prompts, others are negatively affected. A layer-wise analysis indicates that surface-level features such as temporality and negation are captured more reliably than deeper semantic phenomena like quantification. Consistent with prior work, our findings highlight the limited capacity of current LLMs to generate fully formal meaning representations.
Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs ext...
Subavarshana Arumugam, Mamta Nallaretnam, K. Wickramasinghe et al.· 0 citations
A declarative probabilistic framework for pre-training evaluation that makes the semantic structure of model behavior explicit and provides new formal tools for relating evaluation to learning, and links evaluation and learning through a shared semantics.
Kyle Richardson, C. Anderson, Pranav Balakrishnan et al.· 1 citation
AVA is introduced, a systematic framework for evaluating whether embeddings distinguish logic-sensitive relational semantics in ontologies and knowledge graphs, and reveals a persistent gap between linguistic representation learning and ontology-level discrimination, challenging the assumption that strong NLP benchmark...
Hamed Babaei Giglou, Jennifer D'Souza, S. Auer· 0 citations
SLITE is presented, an explainable hybrid model for Recognizing Textual Entailment that integrates two complementary layers of semantic analysis: a structural-relational layer, based on semantic compatibility and incompatibility between compositional entities, and a distributional-informational layer, based on structur...
David Torres-Moreno, J. Hermosillo-Valadez, Asela Reig-Alamillo· 0 citations
While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant inf...
Subavarshana Arumugam, Mamta Nallaretnam, K. Wickramasinghe et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.