Skip to content

From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

Sep 2026 · 1 citation · 50 references
Computer Science

TL;DR

A declarative probabilistic framework for pre-training evaluation that makes the semantic structure of model behavior explicit and provides new formal tools for relating evaluation to learning, and links evaluation and learning through a shared semantics.

Abstract

While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the learning signals used in pre-training. In this paper, we propose ModelLog, a declarative probabilistic framework for pre-training evaluation that makes the semantic structure of model behavior explicit and provides new formal tools for relating evaluation to learning. ModelLog specifies evaluation targets as symbolic constraints over token-level predictions and measures how strongly a model's distribution satisfies those constraints. We explore the framework through a new suite of tasks targeting negation, mutual exclusivity, and consistency, finding systematic failures that are difficult to characterize through token likelihood or answer accuracy alone. We further show that these evaluation scores can also be interpreted as losses, whose gradients reflect logical strength, informativeness, and variable-level sensitivity. This links evaluation and learning through a shared semantics, suggesting evaluation methods that diagnose model behavior while also helping to clarify the semantic structure of learning.

View source

Similar papers

Open access 2026

Semantic Parsing for Evaluating Large Language Models: Separating Linguistic Abilities with YARN

A layer-wise analysis indicates that surface-level features such as temporality and negation are captured more reliably than deeper semantic phenomena like quantification in large language models, highlighting the limited capacity of current LLMs to generate fully formal meaning representations.

Rémi De Vergnette, Maxime Amblard · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs

Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, ra...

Ramya Keerthy Thatikonda, W. Buntine, Ehsan Shareghi · 0 citations
#artificial intelligence Preprint Sep 2026

Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probab...

Fei-Yang Li, Sheng-Jing Liu, Qi Zhan et al. · 0 citations
#natural language process... Preprint Sep 2026

Chronologic: Measuring Language Models'Ability to Represent the Past

Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...

Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al. · 0 citations
Review Open access Sep 2026

Enduring Strengths of Knowledge-Based Methods in the Era of Machine Learning and Large Language Models: A Task-and-Architecture Perspective

Large language models have shifted AI toward statistical learning, but knowledge-based methods remain essential for tasks governed by combinatorial structure, declarative correctness, strong domain priors, and auditable reasoning. This paper treats the issue as one of task–architecture fit rather than paradigm competit...

Maikel Leon · 0 citations
#machine learning Preprint Sep 2026

Evaluation of Contextual Understanding in Large Language Models

Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs ext...

Subavarshana Arumugam, Mamta Nallaretnam, K. Wickramasinghe et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.