ABSTRACT
The intersection of copyright law and artificial intelligence raises fundamental questions about the optimal scope and form of intellectual property protection in a world of generative AI. Drawing on the economics of information goods, the technical literature on memorisation in large language models, and a comparative analysis of recent legal developments across Germany, the United Kingdom and the United States, this paper distinguishes two markets in which AI and copyright interact – the market for training data access and the market for reproduced outputs – and shows that they raise distinct welfare problems. Building on Gans (2024), we propose a capability-based levy that links payments to rights holders to the measurable memorisation rate of deployed models to reflect the potential harm caused in the market for reproduced outputs. We argue that the distribution problem that any such mechanism creates, as well as the non-pecuniary harm to creators from unattributed use of their work, can be addressed by mandating source-attribution capability at the model-architecture level. The capability-based approach has implementation advantages over training-data-based schemes, including reduced exposure to territorial arbitrage and stronger incentives for AI developers to reduce memorisation.
Koboldt, Christian, Copyright and Artificial Intelligence: The Importance of Memorisation and Attribution (May 20, 2026).
Leave a Reply