Token Trajectories as Knowledge Representation in Large Language Models
Abstract
The subject of this paper is the problem of knowledge revealed by the emergence of advanced natural language processing technology, called large language models (LLMs). It proposes an interpretation of the appearance of this phenomenon as an emergent effect of a complex system, such as LLMs. The meaning of this interpretation is based on the observation of the computational trajectory of computational units (vectors) which, on the other hand, represent discrete semantic units, tokens. To justify the autonomy and relevance of this interpretation of knowledge, the paper invokes the concept of discursive space, which allows, in particular, for the reliance on language as a container of knowledge and for the generalization of the idea of knowledge to other semantic environments. To describe more general units of knowledge, the strong version of the theory proposes the introduction of the institution of gnosemes, of which discourse is a special case. Gnosemes find application in the case of LLMs. Due to the different ontical contexts of the entities described, the paper proposes a transdisciplinary approach, the axis of which is the field of social sciences.