Transcription is not generation: distinguishing non-generative AI tool use from academic misconduct in higher education assessment
Abstract
Since 2023, universities worldwide have adopted policies restricting or prohibiting the use of generative artificial intelligence (GenAI) in assessed work. In the United Kingdom, the Quality Assurance Agency (QAA) and the Russell Group have issued guidance, and many institutions now operate tiered assessment classification systems. Internationally, An et al. (2025) report that ninety-four per cent of the top fifty US universities (by the ranking used in that study) had issued faculty-facing GenAI guidelines, and Luo (2024) documents broadly similar structures across the world’ s top twenty institutions by QS ranking. This paper argues that such policies frequently lack the technical precision necessary to distinguish between genuinely generative AI use—where a machine produces the assessed intellectual content—and non-generative uses of AI-powered tools, such as optical character recognition (OCR), voice-to-text transcription, and handwriting recognition, where the tool performs a clerical format-conversion function on content the student has already authored. The paper is a conceptual and policy analysis; it advances interpretive arguments and a proposed evidential framework, not empirical findings. Its principal limitations are stated in Section 7.4. This paper develops a principled framework for distinguishing transcription from generation in the context of academic integrity policy. It examines (i) the technical taxonomy of AI capabilities, drawing on the computer science literature on handwritten text recognition (HTR) and the Organisation for Economic Co-operation and Development (OECD) AI taxonomy; (ii) the purposive interpretation of assessment prohibition categories; (iii) a set of operational criteria that decision-makers could apply to identify transcription-only use; (iv) the reported limitations of stylistic heuristics and AI detection tools; and (v) the implications for assessment validity, procedural fairness, and policy design. On the function-based interpretation defended in this paper, faithful transcription of handwritten material into a typeset format is not best characterised as ‘use of generative AI’ within the purpose of assessment-prohibition policies: the function performed is clerical (converting format without altering content) and does not engage the mischief at which those policies are aimed—the outsourcing of intellectual authorship. With respect to detection tools, the peer-reviewed evidence summarised in Section 5 is mixed, and this paper reports it in both directions. Perkins et al. (2024), evaluating Turnitin’ s AI detection feature on twenty-two experimental GPT-4 submissions marked alongside nine hundred and sixty-three genuine student submissions, report that the tool flagged 91 per cent of experimental submissions as containing some AI-generated content—a result the authors characterise as promising—but identified only 54.8 per cent of the AI-generated content within those submissions, and that faculty, even with the Turnitin score visible, formally reported only 54.5 per cent of the submissions through the misconduct process. Liang et al. (2023) report that for AI-generated US college admission essays a single self-edit prompt dropped detection rates across seven widely-used detectors from up to 100 per cent to up to 13 per cent, and for AI-generated scientific abstracts from up to 68 per cent to up to 28 per cent. Weber-Wulff et al. (2023) conclude that none of the fourteen tools they tested was sufficiently reliable for use as sole evidence. The paper argues that this joint performance—tool plus marker, under adversarial conditions that students are not prevented from using—does not, on the evidence cited, support treating detector output as sufficient for a misconduct finding against the balance-of-probabilities standard many institutions adopt. Bias against non-native English writers is documented by Liang et al. (2023) through the perplexity mechanism; the extension of that mechanism to other populations (for example, writers with structured or systematic prose styles) is argued in this paper as a plausible inference from the same mechanism and is flagged as a hypothesis requiring empirical testing, not as an established empirical finding. The paper proposes four operational criteria—fidelity, non-augmentation, traceability, and attestation—as a structured evidential framework that decision-makers could apply; the framework has not been empirically validated and its inter-rater reliability in practice is an open question. Where assessment policies define the prohibition by reference to generative AI function rather than platform identity, the paper argues that sanctioning students for transcription-only use of AI-powered tools is best characterised as policy misapplication on the interpretation defended here, rather than as a proper finding of academic misconduct. On that interpretation, institutions should consider: defining ‘use of generative AI’ by reference to function rather than platform identity; explicitly distinguishing content generation from format conversion; adopting the operational criteria proposed in Section 4; and requiring validated evidential methods before imposing sanctions. The paper does not claim to have empirically tested these proposals; it sets out an agenda for such testing in Section 7.4.