{"id":20005,"date":"2026-08-11T18:51:47","date_gmt":"2026-08-11T16:51:47","guid":{"rendered":"https:\/\/haimagazine.com\/uncategorized\/paper-for-machines-why-does-ai-need-old-books\/"},"modified":"2026-08-28T11:40:26","modified_gmt":"2026-08-28T09:40:26","slug":"paper-for-machines-why-does-ai-need-old-books","status":"publish","type":"post","link":"https:\/\/haimagazine.com\/en\/hai-premium-2\/paper-for-machines-why-does-ai-need-old-books\/","title":{"rendered":"\ud83d\udd12 Paper for machines. Why does AI need old books?"},"content":{"rendered":"<p class=\"wp-block-paragraph\">A technical book no one has checked out in years. A monograph published in a small print run in the &#8217;80s. A medical textbook that&#8217;s long been out of circulation. Until recently, such books were mostly a problem for libraries and antiquarian bookstores: they took up space, it was hard to find a buyer, and digitization often just wasn&#8217;t worth it.<\/p><p class=\"wp-block-paragraph\">The development of artificial intelligence is changing this landscape.<\/p><p class=\"wp-block-paragraph\">In Poland, the issue has surfaced in recent days after reports about foreign companies interested in buying books from antiquarian bookstores.<mark style=\"background-color:#82D65E\" class=\"has-inline-color has-base-color\"> <a href=\"https:\/\/www.pap.pl\/aktualnosci\/ekspertka-nie-potepiam-antykwariuszy-sprzedajacych-ksiazki-firmom-ai\" target=\"_blank\" rel=\"noopener\">PAP described cases of such offers<\/a><\/mark>, linking the interest in books to the growing demand for data among AI model developers.<\/p><p class=\"wp-block-paragraph\">We also don\u2019t know exactly what happens to every book bought in Poland today or where it ultimately ends up. So there\u2019s no basis for claiming that Polish collections are being shipped en masse to AI labs. We do know, though, that the mechanism itself is perfectly plausible. One of the largest companies developing generative AI has already done something very similar on an industrial scale.<\/p><h4 class=\"wp-block-heading\">Millions of books on the chopping block<\/h4><p class=\"wp-block-paragraph\">We&#8217;re talking about Anthropic, the producer of the Claude models.<\/p><p class=\"wp-block-paragraph\">In the U.S. case <em>Bartz v. Anthropic<\/em>, book authors accused the company of copyright infringement during the creation of datasets used to develop AI models. <a href=\"https:\/\/cand.uscourts.gov\/cases-e-filing\/cases\/324-cv-05417-amo\/bartz-et-al-v-anthropic-pbc\" target=\"_blank\" rel=\"noopener\">The federal court for the Northern District of California provided the public case records<\/a>. The documents indicate that Anthropic built its own digital library in two ways. The company used books previously obtained from unauthorized online libraries, while at the same time it began buying vast quantities of print copies. The court later treated these two practices differently.<\/p><p class=\"wp-block-paragraph\">The process for legally purchased books was surprisingly analog for one of the world&#8217;s most advanced tech businesses. Anthropic spent millions of dollars on millions of copies, including used books. Then external contractors removed the bindings, cut the books apart, scanned the pages, and created digital copies of the text. After that process, the paper copies didn&#8217;t go back on the market.<\/p><p class=\"wp-block-paragraph\">Some of the library built this way was later used to create training datasets for models. <\/p><h4 class=\"wp-block-heading\">The internet isn&#8217;t enough<\/h4><p class=\"wp-block-paragraph\">At first glance, the whole undertaking seems rather absurd. Tech companies have had the internet, digital libraries and massive text corpora for years. Why buy truckloads of paper books, cut them up, and scan them?<\/p><p class=\"wp-block-paragraph\">One obvious answer is this: a large share of human knowledge has never made it onto the internet in a form you can just download and use. This is especially true of older, specialized, local publications and those published in languages other than English. Old textbooks, industry reports, academic dissertations or regional monographs may exist in libraries and private collections while being virtually absent from the open internet.<\/p><p class=\"wp-block-paragraph\">But there\u2019s another problem. The internet itself is changing.<\/p><p class=\"wp-block-paragraph\">The first large language models emerged at a very particular moment. Back then, the vast majority of text online was undeniably human-written. Blogs, forums, Wikipedia, news sites, documents, comments and websites were created before generative AI started producing text at scale.<\/p><p class=\"wp-block-paragraph\">These days, you can&#8217;t make that assumption anymore.<\/p><p class=\"wp-block-paragraph\">Models generate articles, product descriptions, marketing copy, comments, SEO posts and entire websites. So another model that collects data from the web is increasingly running into content that other models created earlier.<\/p><p class=\"wp-block-paragraph\">That doesn&#8217;t mean synthetic data are bad by definition. The problem comes when generated data start replacing real-world data instead of complementing them.<\/p><p class=\"wp-block-paragraph\">In 2024, Ilia Shumailov\u2019s team published <a href=\"https:\/\/www.nature.com\/articles\/s41586-024-07566-y\" target=\"_blank\" rel=\"noopener\"><mark style=\"background-color:#82D65E\" class=\"has-inline-color has-base-color\">a study in &#8220;Nature&#8221; on model collapse<\/mark><\/a>. The researchers examined what happens when successive generations of models are trained on data produced by earlier models. With recursive reuse of such data, the systems gradually lost information about the original distribution, especially about rarer phenomena. In other words, they increasingly generalized away what had originally been the essence of the matter.<\/p><p class=\"wp-block-paragraph\">As more and more content online is generated by models, access to original human-produced data is becoming increasingly important. So a book published thirty or forty years ago now has a quality that no one used to consider particularly valuable: it definitely wasn\u2019t written by a language model.<\/p><h4 class=\"wp-block-heading\">Useless and priceless<\/h4><p class=\"wp-block-paragraph\">This leads to one more change.<\/p><p class=\"wp-block-paragraph\">The book market has always had its own hierarchy of value. What matters is an author&#8217;s popularity, the print run, the edition, the condition of a copy and the level of interest from readers and collectors. For a company building a dataset, those criteria look very different.<\/p><p class=\"wp-block-paragraph\">Imagine a production technology textbook published in the Polish People\u2019s Republic. It doesn\u2019t mean much to the average reader. A collector might walk right past it, too. In a secondhand bookstore, it\u2019d cost a dozen or so zlotys and sit for years waiting for a buyer. Yet it might hold hundreds of pages of specialized language you won\u2019t find on Wikipedia or modern websites.<\/p><p class=\"wp-block-paragraph\">The same goes for older medical, technical, legal and agricultural publications, local historical studies, or books on highly specialized fields. Their value no longer lies in whether someone wants to read them today, but in the fact that they contain text that still isn&#8217;t in the large digital collections.<\/p><p class=\"wp-block-paragraph\">When it comes to Polish, things get even more interesting. There\u2019s far more English-language data online. If you want to expand models\u2019 capabilities in other languages\u2014or in specialized variants of those languages\u2014the pool of digitally available texts starts drying up much faster.<\/p><p class=\"wp-block-paragraph\">Seen this way, an old book isn&#8217;t just a secondhand copy anymore, but a data package with its creation date, language and context.<\/p><h4 class=\"wp-block-heading\">What AI was fed<\/h4><p class=\"wp-block-paragraph\">The timing of when information about book buybacks shows up is interesting for another reason as well.<\/p><p class=\"wp-block-paragraph\">Starting August 2, 2026, the European Commission can enforce the obligations for providers of general-purpose AI models (GPAI) under the AI Act. Those obligations took effect a year earlier, but it&#8217;s only now that the Commission has the enforcement powers that go with them, including the ability to impose fines.<\/p><p class=\"wp-block-paragraph\">One of these obligations is that model developers publish a summary of the content used during training.<\/p><p class=\"wp-block-paragraph\">This isn&#8217;t about disclosing a complete catalog of every text the model has read. <a href=\"https:\/\/digital-strategy.ec.europa.eu\/en\/faqs\/template-general-purpose-ai-model-providers-summarise-their-training-content\" target=\"_blank\" rel=\"noopener\">The Commission\u2019s mandatory template<\/a> requires information about the types of data and their sources. Providers are expected to list, among other things, public and private datasets, content collected from the internet, user data and synthetic data. For web data, they must also include details on the crawlers used, the collection period and the largest sources.<\/p><p class=\"wp-block-paragraph\">One of the goals is to enable copyright holders to better assess whether their content may have been included in the training data.<\/p><p class=\"wp-block-paragraph\">If there&#8217;s a breach of obligations regarding GPAI models, the maximum fine can be \u20ac15 million or 3% of the company&#8217;s global annual turnover, whichever is higher.<\/p><p class=\"wp-block-paragraph\">It&#8217;s a big change, but it doesn&#8217;t solve the whole problem. A book&#8217;s author still won&#8217;t be able to just enter the ISBN and get an answer: yes, this book was included in the training set of a given model.<\/p><p class=\"wp-block-paragraph\">When you buy a book, you&#8217;re buying a specific copy. You can read it, resell it, give it away or throw it out. You don&#8217;t acquire the copyright to the work it contains. In the world of traditional publishing, that was straightforward. Buying a book for a couple of bucks didn&#8217;t entitle anyone to print ten thousand more copies and start selling them.<\/p><p class=\"wp-block-paragraph\">But machine learning doesn&#8217;t fit that old picture.<\/p><p class=\"wp-block-paragraph\">A company can buy a single copy, turn it into data, and then use that data to build a product used by millions of people. There isn\u2019t a traditional copy of the book being offered to customers. Instead, you get a model, and the influence of any particular work within its parameters can be very hard to pinpoint.<\/p><p class=\"wp-block-paragraph\">The debate over how copyright law should handle this kind of process is playing out on both sides of the Atlantic.<\/p><h4 class=\"wp-block-heading\">AI and the offline world<\/h4><p class=\"wp-block-paragraph\">From the start of the digital revolution, paper was supposed to slowly lose relevance. First came search engines, then e-books, large-scale digitization projects and online libraries. Bit by bit, more of our knowledge migrated online.<\/p><p class=\"wp-block-paragraph\">Now it turns out the digital world hasn&#8217;t copied everything.<\/p><p class=\"wp-block-paragraph\">On library shelves, in warehouses, in used bookstores and in private homes, there are still millions of books whose full text has never made it onto the open internet. Some contain outdated knowledge, some are trivial, some are highly specialized. However, from the standpoint of building a massive language corpus, all of them can fill gaps that can\u2019t be filled by scraping the same pages yet again.<\/p><p class=\"wp-block-paragraph\">On top of that, there\u2019s the question of the text\u2019s very origin. The more content models produce, the harder it will be to tell on the internet what\u2019s human-made and what\u2019s already a reprocessing of earlier data by a machine.<\/p><p class=\"wp-block-paragraph\">A 1983 book doesn&#8217;t have this problem.<\/p><p class=\"wp-block-paragraph\">For two decades, we assumed the most valuable resource for internet companies was whatever you could find online. AI is making them increasingly interested in what you can&#8217;t find online.<\/p>","protected":false},"excerpt":{"rendered":"<p>These days, books are gaining a value no one was looking for in them just a few years ago. In a world that&#8217;s increasingly filled with AI-generated content, paper might turn out to be one of the most valuable repositories of texts written exclusively by humans.<\/p>\n","protected":false},"author":465,"featured_media":19869,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"rank_math_lock_modified_date":false,"footnotes":""},"categories":[798,796],"tags":[],"popular":[],"difficulty-level":[38],"ppma_author":[892],"class_list":["post-20005","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-culture-and-media-2","category-hai-premium-2","difficulty-level-medium"],"acf":[],"authors":[{"term_id":892,"user_id":465,"is_guest":0,"slug":"kmironczuk","display_name":"Krzysztof Miro\u0144czuk","avatar_url":{"url":"https:\/\/haimagazine.com\/wp-content\/uploads\/2025\/10\/awatar-2.png","url2x":"https:\/\/haimagazine.com\/wp-content\/uploads\/2025\/10\/awatar-2.png"},"first_name":"Krzysztof","last_name":"Miro\u0144czuk","user_url":"","job_title":"","description":"Od lat zajmuj\u0119 si\u0119 nowymi technologiami w biznesie, edukacji i codziennym \u017cyciu. W centrum mojej uwagi pozostaje cz\u0142owiek \u2013 i to, by technologia wyr\u00f3wnywa\u0142a szanse, zamiast tworzy\u0107 bariery."}],"_links":{"self":[{"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/posts\/20005","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/users\/465"}],"replies":[{"embeddable":true,"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/comments?post=20005"}],"version-history":[{"count":1,"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/posts\/20005\/revisions"}],"predecessor-version":[{"id":20006,"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/posts\/20005\/revisions\/20006"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/media\/19869"}],"wp:attachment":[{"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/media?parent=20005"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/categories?post=20005"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/tags?post=20005"},{"taxonomy":"popular","embeddable":true,"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/popular?post=20005"},{"taxonomy":"difficulty-level","embeddable":true,"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/difficulty-level?post=20005"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/haimagazine.com\/en\/wp-json\/wp\/v2\/ppma_author?post=20005"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}