STEMMA is introduced, a multi-modal and multi-agent framework in which role specific agents collaboratively probe self identification behavior in different models to address concerns about output homogeneity, model biases, and accountability.
Abstract
Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller student model while significantly reducing computational costs. However, as the use of distillation increases in both scale and complexity it raises an important question about what kind of knowledge is really transferred from the teacher model. In this work, we argue that apart from the functional knowledge, student models also learn behavioral patterns, specifically how a model represents its own identity raising concerns about output homogeneity, model biases, and accountability. To address this challenge, we introduce STEMMA, a multi-modal and multi-agent framework in which role specific agents collaboratively probe self identification behavior in different models. We also contribute a set of adversarial prompts designed manually to evaluate identity consistency in LLMs. Our results show that to an extent most models are vulnerable to inconsistencies in self-representations.
It is suggested that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs...
Rong-Can Pei, Zhepei Wei, Shu-Yao Xu et al.· 1 citation
Resistance, the rate at which a model keeps its correct answer under this pressure, is matched with Receptivity, the rate at which a model adopts a correct peer answer after initially answering incorrectly, and six methods are scored.
This position paper observes that a wide variety of techniques designed to improve specific aspects of LM behavior-targeting properties as diverse as adversarial robustness and factual coherence-can be understood as special cases of a common "consistency optimization" procedure and addressed with a standard set of opti...
Itamar Hagay Pres, Belinda Z. Li, L. Ruis et al.· 2 citations
A TBC-TBA self-refine learning multi-agent framework that enables dynamic normative adaptation through iterative multi-agent feedback that integrates Think-Before-Chat (social feedback processing) and Think-Before-Act (norm-guided decision making) phases, allowing agents to progressively refine their normative understa...
Rong-Xin Cheng, Jianhui Yang, Bo-Han Xiong et al.· Proceedings of the Thirty-Fi...· 0 citations
This work proposes a refined security principle: only agents capable of deliberative System 2 reasoning may access untrusted documents, and empirically compares state-of-the-art reasoning language models with standard language models to show that reasoning-capable models are substantially more robust to corrupted evide...
Mehrdad Ghassabi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.