posted in Technology

AI labs buy, scan, shred millions of rare books

Artificial intelligence labs are in a new arms race to buy up millions of rare books, slicing them open, scanning the pages and pulping the remains — sparking concerns that the last remaining copies of out-of-print texts are being destroyed on an industrial scale.

ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.

“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”

www.news.com.au/technology/online/internet/ai-labs-buy-scan-shred-millions-of-rare-books/news-story/0c3b45a67093ab462a587a0348538ce9

Replying to @⁨themachinestops@lemmy.dbzer0.com⁩

Does anyone remember circa 2014 that radical jihadists in the middle east were destroying ancient and medieval artifacts, including stone heads? And I remember here in the liberal west we thought that was awful.

This is not the only thing that makes me think of that. Also Trump’s recent calls to reduce the size of US national parks and open them up to oil and uranium mining exploitation.

Replying to @⁨uriel238@lemmy.blahaj.zone⁩

Yeah we remember. Because that’s a literal fucking warcrime. You’re not allowed to destroy historical sites.

While I agree that destroying books is bad - even common ones - it’s also not technically illegal. If I want to burn a first edition of whatever book worth thousands of dollars… I’m legally allowed to do that.

I can also see the benefit of having a digital version of a rare book that can be enjoyed by thousands as opposed to sitting on a shelf. But yeah, definitely hold companies accountable for destroying them. There has to be a better way to do that work and still have a useable book afterwards.

Replying to @⁨FinishingDutch@lemmy.world⁩

If they were archiving digital versions so they could be read by thousands in the future, this would be a totally different conversation.

They don’t even need to be as selfless as the internet archive - just follow Googles lead and allow people to search and view selections the archive, with automatic opening of the archive when you are confident copyright is expired. Even a closed archive owned by a third party with a dead mans switch to open it in the future when the business model is done would be something.

en