Skip to content

Author

Fabio De Ponte

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

Population-level linguistic conventionality benefits the robustness, learnability, and cognitive efficiency of emergent languages

Human languages are widely understood to be conventional: for a given meaning, a population of language users expect a certain form to be used. Given the lack of unconventional natural languages with which to compare, there is little direct empirical evidence which explains why linguistic conventionality matters. We carry out a series of emergent language experiments with artificial agents designed to exhibit the effect of conventionality on robustness, learnability, and cognitive efficiency. By experimentally manipulating the interaction dynamics through which agents align their vocabularies, we create conditions that either promote or prevent population-level convergence, resulting in conventional and unconventional emergent languages while holding communicative success constant. We find that conventional languages are indeed more robust, better learnable, and more efficient, when compared with less conventional languages.

Jamie D Wright, Lara Verheyen, Fabio De Ponte et al. · 0 citations
#small language model Preprint Aug 2026

Jokes Aside: Measuring the Semantic Distance of Double Meanings

Large language models have significantly enriched the toolkit for computational humor research, particularly in the automated generation of jokes and puns. A key innovation, contextual embedding vectors, offers new opportunities to revisit and refine earlier hypotheses. Notably, Petrovic and Matthews (2013) proposed a joke generation model based on the scheme"I like my X like I like my Y, Z"(e.g."I like my ice like I like my dreams, crushed"). They suggested that joke hilarity increases with: a) frequent association of Z with X and Y, b) rarity of Z, c) ambiguity of Z, and d) meaning distance between X and Y. Building on this, Winters et al. (2019) proposed a set of metrics, based on Google Ngrams and Word2Vector. In this work, three out of their five metrics are revisited with word embeddings: obviousness, compatibility, and comparison. Another measure, symmetry, defined as closeness of Z to both X and Y, is introduced here for the first time. Two models were used to collect the embedding vectors (OpenAI text-embedding-3-small and MiniLM all-MiniLM-L6-v2) on three datasets: JokeJudger, Expunations, and rJokes. The last two datasets, Expunations, and rJokes, were expanded by adding paired sentences that captured the ambiguous expression at the core of each joke in its two different meanings. Results revealed that models trained on the proposed metrics performed poorly in predicting humor ratings: on JokeJudger, the best model achieved 57.1% accuracy, below the 61.5% baseline, while performance on Expunations and rJokes was even lower. Nevertheless, the symmetry metric seems consistently associated with higher-rated jokes, suggesting it may capture a necessary -though not sufficient- property of humor.

Fabio De Ponte · 0 citations
Preprint Jul 2026

Evolving language compositionality in a frequency-structured meaning space

The iterated learning model was introduced to investigate language evolution: the way in which the characteristic properties of human languages have been shaped, at least partly, by repeated transmission from one language user to another. The key finding is that language compositionality can arise spontaneously as a consequence of language being passed repeatedly through a language learning bottleneck. Here we explore how changing the frequency of different meanings, so that some meanings occur much more frequently than others, affects the character of its compositionality. We find that, as observed in natural languages, high-frequency meanings can escape the pressure to conform to the grammar that characterizes lower-frequency meanings. However, when the frequency structure is instead imposed on parts rather than on whole meaning vectors, the language fails to transmit across generations. This occurs despite the fact that the most frequent elements are reliably learned. These results suggest that frequency can shape emergent linguistic structure only when the frequency distribution is defined over form-meaning units that learners can acquire holistically. When frequency is instead distributed over smaller units, it fails to support the relational structure required for compositional generalisation, thereby preventing stable language transmission.

Fabio De Ponte, Eloise Gaines-White, Conor J. Houghton et al. · 0 citations