Skip to content
Preprint

STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

STEMMA is introduced, a multi-modal and multi-agent framework in which role specific agents collaboratively probe self identification behavior in different models to address concerns about output homogeneity, model biases, and accountability.

Abstract

Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller student model while significantly reducing computational costs. However, as the use of distillation increases in both scale and complexity it raises an important question about what kind of knowledge is really transferred from the teacher model. In this work, we argue that apart from the functional knowledge, student models also learn behavioral patterns, specifically how a model represents its own identity raising concerns about output homogeneity, model biases, and accountability. To address this challenge, we introduce STEMMA, a multi-modal and multi-agent framework in which role specific agents collaboratively probe self identification behavior in different models. We also contribute a set of adversarial prompts designed manually to evaluate identity consistency in LLMs. Our results show that to an extent most models are vulnerable to inconsistencies in self-representations.

View source

Similar papers

Jul 2026

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

It is suggested that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.

Marylou Fauchard, Florian Carichon, Margarida Carvalho et al. · 0 citations
#machine learning Preprint Sep 2026

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs...

Rong-Can Pei, Zhepei Wei, Shu-Yao Xu et al. · 1 citation
Preprint Jul 2026

Position: It's Time to Optimize LLMs for Self-Consistency

This position paper observes that a wide variety of techniques designed to improve specific aspects of LM behavior-targeting properties as diverse as adversarial robustness and factual coherence-can be understood as special cases of a common "consistency optimization" procedure and addressed with a standard set of opti...

Itamar Hagay Pres, Belinda Z. Li, L. Ruis et al. · 2 citations
Conference Open access Sep 2026

Self-Refine Learning in LLM Multi-Agent Systems for Legal Norm Cognition and Compliance

A TBC-TBA self-refine learning multi-agent framework that enables dynamic normative adaptation through iterative multi-agent feedback that integrates Think-Before-Chat (social feedback processing) and Think-Before-Act (norm-guided decision making) phases, allowing agents to progressively refine their normative understa...

Rong-Xin Cheng, Jianhui Yang, Bo-Han Xiong et al. · 0 citations
#natural language process... Preprint Aug 2026

Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents

This work proposes a refined security principle: only agents capable of deliberative System 2 reasoning may access untrusted documents, and empirically compares state-of-the-art reasoning language models with standard language models to show that reasoning-capable models are substantially more robust to corrupted evide...

Mehrdad Ghassabi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.