What a nano-GPT can (not) tell us about spoken language
Abstract
After the launch of ChatGPT in autumn 2022, a lot of research has focussed on the quality and near-naturalness Large Language Model-based tools present in the texts they produce. While one area of research has focussed on the similarities and differences between machine-produced and human-produced output (e.g. Berber Sardinha, 2024 ), others have explored how far such tools could process more complex tasks (e.g. Curry et al., 2024 ; Valmeekam, et al., 2023 ). While it can be assumed that ChatGPT makes use of written-to-be-spoken training material, there has been no investigation, as yet, into how far a Generative Pre-trained Transformer (GPT) algorithm is able to process (transcribed) natural, colloquial language. This research will investigate whether spoken language transcripts lead to processing difficulties; whether such generated language can be seen as a suitable reflection of natural speech; and whether machine produced texts offer new insights into the workings of language.