Skip to content
Open access

Semantic Parsing for Evaluating Large Language Models: Separating Linguistic Abilities with YARN

2026 · International Conference on Language Resources and Evaluation · pp. 9745-9755 · 0 citations · 28 references
Computer Science

TL;DR

A layer-wise analysis indicates that surface-level features such as temporality and negation are captured more reliably than deeper semantic phenomena like quantification in large language models, highlighting the limited capacity of current LLMs to generate fully formal meaning representations.

Abstract

We evaluate large language models (LLMs) through semantic parsing into Yarn, a structured meaning representation that distinguishes predicate–argument structure from higher-level linguistic features such as tense, aspect, and modality. For evaluation, we employ SmatchY, a fine-grained metric designed to assess different layers of meaning independently. Our experiments test multiple LLMs under varied conditions, including inference modes, linearization formats (JSON and logic-inspired CFG), and the presence or absence of auxiliary supervision via partial semantic parses. Results show that model performance is highly sensitive to both representational design and supervision, with no single configuration consistently outperforming the others. While some models gain from additional semantic information in prompts, others are negatively affected. A layer-wise analysis indicates that surface-level features such as temporality and negation are captured more reliably than deeper semantic phenomena like quantification. Consistent with prior work, our findings highlight the limited capacity of current LLMs to generate fully formal meaning representations.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Evaluation of Contextual Understanding in Large Language Models

Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs ext...

Subavarshana Arumugam, Mamta Nallaretnam, K. Wickramasinghe et al. · 0 citations
#natural language process... Preprint Sep 2026

From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

A declarative probabilistic framework for pre-training evaluation that makes the semantic structure of model behavior explicit and provides new formal tools for relating evaluation to learning, and links evaluation and learning through a shared semantics.

Kyle Richardson, C. Anderson, Pranav Balakrishnan et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Do General NLP Embeddings Capture Ontological Reasoning?

AVA is introduced, a systematic framework for evaluating whether embeddings distinguish logic-sensitive relational semantics in ontologies and knowledge graphs, and reveals a persistent gap between linguistic representation learning and ontology-level discrimination, challenging the assumption that strong NLP benchmark...

Hamed Babaei Giglou, Jennifer D'Souza, S. Auer · 0 citations
#natural language process... Preprint Sep 2026

Linguistic Features for Interpretable Textual Entailment

SLITE is presented, an explainable hybrid model for Recognizing Textual Entailment that integrates two complementary layers of semantic analysis: a structural-relational layer, based on semantic compatibility and incompatibility between compositional entities, and a distributional-informational layer, based on structur...

David Torres-Moreno, J. Hermosillo-Valadez, Asela Reig-Alamillo · 0 citations
#artificial intelligence Preprint Sep 2026

Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework

While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant inf...

Subavarshana Arumugam, Mamta Nallaretnam, K. Wickramasinghe et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.