ABSTRACT
Generative AI models are trained on massive datasets scraped from the web, but many of these data points are protected by copyright. This raises a fundamental tension: how can we build AI that is both legally compliant and unbiased? We analyze LAION-5B, one of the world’s largest public image datasets, to measure the legal status of its content and assess how copyright-based filtering affects technical and semantic representation. We find that fewer than 0.01% of images are clearly licensed for unrestricted use. Filtering by license information not only shrinks the dataset dramatically but also introduces severe bias. We then show that an ‘opt-out’ system, where creators can exclude their work, can be more practical and less distorting than requiring advance permission. Our findings expose a regulatory paradox at the core of AI governance-compliance with copyright law can conflict with the representativeness that emerging AI regulations demand. We discuss policy alternatives, including data markets, statutory licensing, and safe harbor provisions, that could help reconcile these competing goals and build legal infrastructure suited for the scale of modern AI.
Shcherbakov, Viktor and Dalaud, Irvin and Peukert, Christian, AI needs better data than the law allows (March 14, 2025).
Leave a Reply