Copyright Infringement Defenses for Generative AI Training: Comparing U.S. Fair Use, EU TDM Exceptions and the AI Act
Abstract
Training generative artificial intelligence models often involves extensive reproduction of copyrighted works. Whether such training may lawfully be conducted without authorization has consequently become a major point of contention in contemporary copyright law. This article systematically examines the legal characterization of the use of copyrighted works during the training of generative AI models, focusing on two central questions: whether reproductions made during AI training fall within the scope of copyright infringement and whether existing defenses, particularly U.S. fair use and EU text and data mining (TDM) exceptions, can justify such conduct. Through a comparative analysis of the literature and recent developments including Andy Warhol Foundation v. Goldsmith and Kneschke v. LAION, this article argues that both systems face structural limitations. In the United States, the increasingly restrictive interpretation of transformative use and growing concern over market substitution weaken the certainty of fair‑use defenses. In the European Union, the practical operation of the opt‑out mechanism and transparency obligations remains problematic. The phenomenon of model memorization further challenges the theoretical foundation of non‑expressive use. The article concludes by evaluating collective licensing, statutory licensing, taxation mechanisms, and technological safeguards as possible reform directions.