Skip to content

The UK needs more transparency from AI companies digitising books to train models

AI companies using physical sources of media, like books, to train models skirt existing data laws. We need more transparency.

The UK needs more transparency from AI companies digitising books to train models
Photo by Gülfer ERGİN / Unsplash

In their hunger for fresh training data, AI companies are creating an "anti-library:" buying second-hand books at an industrial scale, scanning them and destroying the originals. Nobody knows how many books are disappearing from circulation in the process. Yet the government’s recent call for evidence on Data Regulation in the Age of AI overlooks this practice. 

The Department for Science, Innovation & Technology, which led the consultation, asked important questions about the provenance, traceability and accountability of data that organisations want to access or share, but not what happens when companies acquire physical sources, digitise and destroy them in the process. The consultation asks how data should be governed across global, opaque supply chains involving several organisations, which is a fair description of the second-hand book pipeline.

Yet its question about “how roles and responsibilities are determined across organisations and throughout the data lifecycle” has a simple answer in this case: they aren’t. 

Why books, and why now

To answer this, we need to look back at the last decade of AI innovation.

This content is for members only

Subscribe
Add The Stack on Google