AI Training & Copyright: Between Black Box and Legal Framework
Key Takeaways
- For the first time in Europe, a court has held that training generative AI models on protected works without consent constitutes copyright infringement.
- The Munich Regional Court (LG München I) treats the memorization of training data in a model's parameters as a reproduction under Section 16 of the German Copyright Act (UrhG), despite the technical black-box problem and a widely held opposing view in computer science.
- Independently of that, the output of protected works to users infringes copyright in its own right (Sections 16, 19a UrhG). The court did, however, dismiss a personality-rights claim over distorted outputs, and it left open the question of distortion under Section 14 UrhG.
- In the court's view, neither Section 44a UrhG nor the text-and-data-mining exception in Section 44b UrhG covers the training itself.
- German copyright law also applies to US providers when outputs are intentionally delivered to users in Germany.
Training AI systems on copyrighted works
AI systems are trained on vast datasets that often include copyrighted works. That is what enables them to generate text, images, and other creative outputs on their own in response to a prompt. It also raises an immediate question: To what extent is training generative AI models on protected works permissible under copyright law? The legal analysis is complicated by the fact that AI systems are a so-called "black box": their inner workings can be traced only to a limited degree. Against this backdrop, the Munich Regional Court (LG München I) became one of the first courts in Europe to address whether the memorization of copyrighted works in generative language models constitutes a reproduction, in the dispute between GEMA and OpenAI (judgment of the LG München I of November 11, 2025, case no. 42 O 14139/24).
Starting from that landmark ruling, this article offers an accessible overview of the current debate around the copyright status of AI training.
GEMA v. OpenAI
GEMA, the German music rights collecting society, sued the US company OpenAI for infringing copyrights in song lyrics, specifically over the training of ChatGPT models on license-protected lyrics without GEMA's consent. The court concluded that, in this specific case, both the training of ChatGPT on the lyrics at issue and their output to users infringed copyright. What makes the ruling especially interesting are the standards the court applied to get there.
"Memorization" through training
A pivotal part of the judgment concerns the hotly debated phenomenon of memorization of protected works through AI training. Memorization describes the fact that training data can end up contained in AI models and be reproduced identically or near-identically in outputs through a simple prompt. That memorization can occur through the training of AI models (as it did in GEMA v. OpenAI) is undisputed.
Does "memorization" equal reproduction under copyright law?
What is disputed is how exactly memorization arises, what form it takes, and whether, depending on the answer, it can constitute a reproduction under copyright law. A reproduction within the meaning of Section 16 UrhG is any physical fixation of a work capable of making the work perceptible to the human senses, directly or indirectly. Because AI systems are a black box, different technical perspectives lead to different answers on whether memorization amounts to reproduction. In GEMA v. OpenAI, the court concluded that the memorization of the song lyrics in the models is a reproduction within the meaning of Section 16 UrhG. It grounded the physical fixation of the lyrics in the model parameters in the memorization itself: The lyrics are fixed, stored, contained in the models because they can be reproduced in response to a simple prompt. In the court's view, it also suffices for indirect perceptibility that the lyrics can be made perceptible at the output level through simple prompts. The court's reasoning is heavily value-driven: It relies chiefly on the technological neutrality of Section 16 UrhG (paras. 177 et seq.) and on parallels to accepted forms of reproduction such as lossy MP3 compression and progressively stored JPEG files (paras. 183 et seq.). A concretely delineable dataset within the model is, in the court's view, not required (paras. 184 et seq.).
The real doctrinal question is therefore less whether the model contains a copy in the classic sense (the court expressly leaves open whether one speaks of storing, copying, or a mere reflection in parameters, para. 186) than whether the concept of physical fixation extends to the distribution of a work across probability parameters. Computer science scholarship argues that while verbatim reproductions can occur, they are rare and unintended, and it does not follow that entire works are stored in the model as files or text fragments. The reason lies in how AI models technically work: A model's parameters encode nothing more than statistical relationships, and outputs vary within a high-dimensional parameter space. They are not predictable. If an output nonetheless reproduces training data almost identically, on this view that is the result of statistical clustering in the training data (data that appears multiple times, for example), not evidence of a stored copy. Whether memorization therefore falls outside the concept of reproduction under Section 16 UrhG depends decisively on how broadly one construes physical fixation, and the Munich court construes it deliberately broadly.
Output as an independent copyright infringement
The court examines memorization and the actual output of the lyrics as two separate levels of infringement. Even if one does not classify memorization within the model as a reproduction, the output of a protected work to the user is, in itself, a reproduction under Section 16 UrhG and a communication to the public under Section 19a UrhG. Where the model outputs lyrics identically or near-identically in response to a simple prompt, the court sees a direct act of use by the operator. The court did, however, dismiss the claim, based on general personality rights, to stop attributing the distorted outputs to the lyricists (paras. 301 et seq.): Only the authors' social sphere was affected, and there were no serious adverse consequences. It expressly left open the question of distortion under Section 14 UrhG (para. 292), since the injunction claim already followed from Sections 16, 19a, and 23 UrhG.
Other key points of the ruling
At least as significant for practice are two further course-setting determinations:
Text-and-data-mining exception (Section 44b UrhG): The court distinguishes between assembling the training corpus (covered by Section 44b) and training the model itself. Reproductions that arise during training through memorization no longer serve the purpose of text and data mining and are therefore not covered by the exception (paras. 193, 200 et seq.). For commercial providers, that leaves only licensing or training on an effectively deduplicated corpus; the research exception in Section 60d UrhG is available only to non-commercial research institutions.
Liability for outputs: Under the ruling, operators control the acts of use for outputs generated in response to simple prompts (paras. 275 et seq.). That control can shift to the user, however, where outputs are deliberately provoked through manipulative prompts, a distinction that matters greatly in practice but is delicate to draw in individual cases.
Section 44a UrhG (temporary reproductions): The court declines to privilege memorization under Section 44a UrhG. In its view, the fixation in the model parameters is not temporary but permanent: it persists for the model's entire service life and is constitutive of its function.
Opt-out under Section 44b(3) UrhG: The court did not have to rule on the validity of the rights reservation GEMA had declared, since it considered the TDM exception inapplicable to training anyway (paras. 193, 210). In practice, the machine-readable rights reservation is nonetheless likely to remain rights holders' most important lever for prohibiting text and data mining and, on the court's reading, the training built on top of it.
OpenAI as proper defendant: The court affirms both the applicability of German copyright law and the standing of the US OpenAI entities as defendants. What matters is that the outputs at issue are intentionally delivered to users in Germany; separating US-based training from the German output market does not help the provider.
How exactly memorization arises and takes shape, and whether (and if so, in which constellations) it can constitute a reproduction under copyright law, remains unsettled. From a practical perspective, the core question is whether AI training qualifies as a use that requires a license, which would have far-reaching consequences: new licensing structures and the revenue streams that come with them. The answers depend largely on penetrating the technical complexity of AI models. Neither the courts nor the literature have yet developed a consistent line. What is certain is that, building on the Munich ruling, which is not yet final, other trial courts will take up these questions, and further judgments outside this proceeding, particularly at the international level, are eagerly awaited.



