Emergence of Social Reasoning through Adaptive Evolution of LLMs
Abstract
Large language models (LLMs) exhibit advanced social reasoning capabilities like Theory of Mind (ToM), yet the dynamic acquisition process remains underexplored. We propose a constructive, artificial-life approach that treats LLMs as “model organisms,†investigating the emergent mechanisms of cognitive functions through biological adaptive evolution rather than static analysis, as a step toward understanding the human-AI societies now taking form. By applying a genetic algorithm to evolve LoRA adapters from a behaviorally degraded state, in which task performance is reduced to near-random levels while latent knowledge remains in the frozen weights, we analyzed this evolutionary process at both behavioral and mechanistic levels. At the behavioral level, comparing adaptations in knowledge-intensive (MMLU) and social reasoning (ToMBench) environments revealed that task characteristics dictated fitness landscape ruggedness. An asymmetric generalization was observed: while moderate adaptation to broad knowledge partially bolstered heuristic social reasoning, excessive specialization created an evolutionary trade-off constraining deep inferential capabilities. At the mechanistic level, a Sparse Autoencoder (SAE) revealed the dynamic refinement of reasoning mechanisms during ToM evolution. The evolved individual’s strategy underwent a stepwise transition from superficial linguistic cues to mental state concepts, ultimately specializing in ToM-related conceptual representations. This stepwise acquisition trajectory, alongside the compensatory reasoning observed in the knowledge-intensive environment, suggests a structural generality in the adaptive acquisition of higher-order cognitive capabilities. Data/Code available at: https://doi.org/10.5281/zenodo.20790937