AI & Tech 4 min read

AI Companies Are Buying Books They Cannot Replace

AI companies are reportedly turning to shelves of older books for cleaner training data. The story is bigger than one strange bulk order: it reveals what the internet loses when speed becomes the only goal.

A

Author

Writer

An AI scanning machine extracts glowing data from an open old book beside a stack of human-authored books

Table of Contents

    AI companies are buying books in unusually large quantities, and a strange request has been making its way through the world of old books: thousands of copies, sometimes far more, bought in bulk by people who do not always explain what they need them for. The trail has led back to the AI boom.

    Recent reporting from 404 Media describes companies offering to acquire printed books for AI training, presenting them as a source of edited, human-made material that has not already been polluted by generated text. Follow-up reports say booksellers in Europe have received unusually large orders, raising worries about books being scanned, stripped apart and discarded in the process.

    It is a story with enough irony to make anyone stop scrolling. The technology sold as a shortcut to writing, research and creativity is now going back to physical books because the open web is becoming less trustworthy as a source of language.

    The real shortage is not data

    There is no shortage of words online. There are more words than any person could read in a lifetime. What is getting harder to find is work that has passed through a human mind with care: an argument that was checked, a paragraph that was rewritten, a strange detail kept because it mattered rather than because it filled space.

    Books are valuable to an AI company for obvious reasons. They are usually edited. They have structure. They carry a voice from beginning to end. Many were written before generative AI could flood search results, social feeds and self-publishing platforms with endless variations of the same thin idea.

    That last part matters. A language model trained repeatedly on generated material risks learning from a photocopy of a photocopy. The original edges soften. The mistakes multiply. The prose begins to sound smooth but oddly empty. Anyone who spends time online has felt it: five posts saying nearly the same thing, all with the same polished confidence and none with much to say.

    A stack of worn old books beside a translucent sheet representing machine-readable data
    Older books offer something the open web increasingly struggles to provide: work shaped by human editing and judgment.

    Destructive scanning turns a cultural question into a practical one

    Not every used book is rare, irreplaceable or destined for a library shelf. Bookshops routinely deal with surplus stock, damaged copies and titles that have been out of demand for years. That deserves to be said plainly, because panic is not analysis.

    But volume changes the question. When huge orders move through a market built on odd, overlooked and out-of-print books, it becomes difficult to shrug it off as ordinary resale. A book that is easy to replace today may not be easy to replace after a few years of bulk buying, especially when nobody knows who is acquiring it or why.

    The reporting does not prove that every bulk order came from an AI company, nor that every book was destroyed. It does show why people are uneasy. The process is opaque, and opacity is exactly where public trust goes to die.

    What this says about the AI race

    The industry often talks about data as though it were oil: a raw material waiting to be extracted. Books make that metaphor look cheap. A good book is not a barrel of language. It is years of reading, reporting, memory, judgement, failure and revision held together by one person or a small group of people.

    That does not mean AI should never learn from books. Digitisation has real value. Searchable archives, accessible text and preservation copies can expand access to knowledge. Libraries and archives have shown that careful scanning can serve the public.

    The difference is purpose and responsibility. Preservation keeps a work available. A private pipeline that treats books as disposable feedstock may make a model stronger while leaving readers, writers and booksellers with less clarity than before.

    My view

    I do not think the answer is to romanticise every printed book or pretend that new technology is automatically destructive. That would be too easy. But I do think the AI industry is revealing an uncomfortable truth: human work is still the thing it needs most, even as it tries to convince us that human work is becoming optional.

    If companies want to build systems on the accumulated work of writers, researchers and publishers, they should be open about where that material comes from, how it is acquired, and what happens to the physical books involved. “The optics problem” is not a public-relations inconvenience. It is a signal that the public has a reasonable question.

    Conclusion

    The most interesting part of this story is not that machines are learning from books. Machines have always learned from human records in one form or another. The interesting part is that, in an internet increasingly filled with automated noise, old books have become valuable again because they still contain something the noise does not: deliberate human thought.

    That should be a warning to the AI industry, but also a reminder to readers. The work worth keeping is rarely the work produced fastest. It is the work someone cared enough to make specific, honest and lasting.

    Reporting referenced in this article: 404 Media, Tom’s Hardware and NL Times.

    A

    Author

    Writer at Blogsloop

    Author at Blogsloop.