Vue lecture

AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop

AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop

As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing. 

“The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.” 

In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models. 

“Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...]  “Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.”

ISBN stands for International Standard Book Number, the numerical commercial book identifier and barcode on the back of most books. For years, ISBNdb helped book sellers, libraries, and distributors manage their inventory and find and sell books, but the generative AI boom has made it valuable to AI companies. In addition to selling access to book metadata, ISBNdb now helps AI labs source bulk printed book purchases of between 1,000 to 1 million books per order. ISBNdb’s data makes it easier for AI companies to methodically acquire, scan, and turn printed books into training data while avoiding duplication. 

AI companies’ attempts to hoover up printed books for training data got wide attention in January after a copyright lawsuit from book authors against Anthropic revealed internal documents detailing its plan to obtain and scan millions of printed books, and destroy them in the process. The Washington Post article found that Anthropic was buying books from one company called Better World Books, one of several marketplaces where libraries, retailers, and individuals can sell their books. Google was recently sued by book publishers for similarly training Google Gemini on copyrighted books.  

ISBNdb advertises that it can keep the identity of AI companies secret. 

“Strict NDA [non-disclosure agreement] on every engagement,” ISBNdb’s site says. “Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed.”

ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process. 

“The optics problem is real,” ISBNdb’s site says. “‘AI company destroys two million books’ is not a headline that generates sympathy.”

One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. This bookseller asked to remain anonymous so he can continue to do business on these platforms. 

“I personally have mixed feelings about all of this,” the bookseller, who suspects he’s sold hundreds of books to AI companies for training data, told me. “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”

This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain. 

The seller told me that, normally, on a good week, he’d sell about 20 books. Since April, he has regularly sold hundreds of books a week. While the seller didn’t have clear evidence that the purchases were being made by AI companies, the purchases made him suspect that they were. First of all, he said, the kind of books he sells are specialized and are usually bought by schools and libraries. Purchases from these organizations have been trending downward because of reduced funding, he said. Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases. I have not seen any evidence that this bookseller’s recent sales were facilitated by ISBNdb or that the client was an Anthropic or another AI company.  

“It's not just the quantity, but the weirdness of the orders,” the bookseller told me. “I've had library orders before, and usually they're mostly confined to a single subject or maybe a slightly broader range of subjects. But basically, almost every library in the world has lost their budget. I know all the U.S. college libraries don't buy much anymore. The Australian libraries don't buy much anymore. The type of books [...] there's no rhyme or reason to it. Also, there's a total disregard for the price of the book. I've had some books that sold through this way that were [...] greatly overpriced. That's kind of a tell for AI because they have just so much money.” 

“Is it just me, or has there been an uptick in the number of AutoBuy orders since the tail end of last year?” one bookseller wrote on the forums for Alibris, another marketplace for selling books, in February. The AutoBuy function allows a customer to flag books they want to automatically purchase once they become available for sale on Alibris. “Any comment on what is happening? Is an AI going to read every single book? Any insight into how the selections are made? They seem to vary quite a bit in condition, format (hardcover and softcover), price and so on.”

“We have a couple of new bulk buyers that are scooping up trade books so lots of sellers are getting lots of orders,” Mike Feldman, director of client services at Alibris, responded. 

One bookseller told me that similarly large orders of books were coming through another marketplace called Biblio. Customers can provide Biblio with a spreadsheet of ISBNs they want to purchase and the company takes it from there. 

In June, a publication in the Netherlands talked to several rare booksellers who reported similar large bulk purchases they assumed were coming from AI companies. 

It’s hard to say for a fact that the books are being bought for training data and possibly being destroyed by AI companies because ISBNdb and book marketplaces like Biblio and Alibris keep the identity of the buyer hidden. Large bulk purchases of books are first sent to distribution centers where, for example, Alibris checks the quality of the books before sending them off to the client.  

Internal Anthropic documents about its plan to scan millions of books, revealed in the copyright lawsuit, don’t make clear why the company wanted to destroy the books in the process. A deposition of Tom Harvey, who Anthropic hired to lead the project and who previously helped create Google Books, shows that one company Anthropic contracted to scan the books was Datamation, which offers both “high volume destructive and non-destructive book scanning” services. In a destructive book scanning process, the spine of the book is cut so the pages can be fed into a scanning machine, which is faster and cheaper than non-destructive book scanning.

Regardless of its original intentions, the federal judge in the copyright lawsuit from authors against Anthropic, William Alsup, found that Anthropic’s creation of digital copies of the books was legal specifically because the books were destroyed. 

“Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy,” Alsup wrote in his ruling. “The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company.” 

This, Alsup said, was “clearly transformative” and therefore qualified as fair use under Section 107 of the Copyright Act. 

ISBNdb’s site advertises this legal argument to AI companies as well. 

“Purchasing paper books at scale from the secondary market does not deprive any creator of income they would otherwise have received,” ISBNdb’s site says. “These are books that have already fully discharged their financial obligation to their creators.”

“Responsible physical sourcing is not book burning,” ISBNdb’s site in a section about why it’s crucial to recycle the destroyed books. “It is the completion of a book’s lifecycle: from tree to knowledge to tree again.”

ISBNdb and Anthropic did not respond to a request for comment.

  •  

Is the Best Game of the Year a Failure? (With Rob Zacny)

Is the Best Game of the Year a Failure? (With Rob Zacny)

If you listen to the 404 Media podcast by now you probably realized that Joe and I are a little obsessed with a game called Marathon. I’m embarrassed to say that I’ve played it for almost 300 hours since it was released in March.

But as much as we’re enjoying it, and there are thousands of players who feel the same, Marathon so far has failed to find the audience we’d expect from the developer that made Halo, Destiny, and which a few years ago acquired by PlayStation for more than $3 billion. It’s bad news for Marathon fans and a good sign for how much the video game business has changed over the years.  

I wanted to have Remap Radio host Robert Zacny on the podcast because much like me and Joe, he’s been obsessed with Marathon as well. One of Rob’s greatest skills is dissecting how and why games get their hooks into us, and what a game’s popularity, or lack thereof in Marathon’s case, might reveal about the state of the industry and culture more broadly. 

404 Media is a journalist-founded company and needs your support. To subscribe, go to 404media.co. As well as bonus content every single week, subscribers get access to additional episodes where we respond to their best comments. Subscribers also get early access to our interview series. Gain access to that content at 404media.co.

Listen to the weekly podcast on Apple Podcasts, Spotify, or YouTube

Become a paid subscriber for early access to these interview episodes and to power our journalism. If you become a paid subscriber, check your inbox for an email from our podcast host Transistor for a link to the subscribers-only version! You can also add that subscribers feed to your podcast app of choice and never miss an episode that way. The email should also contain the subscribers-only unlisted YouTube link for the extended video version too. It will also be in the show notes in your podcast player.

  •  

Scammers Sell Seeds for Exotic AI-Generated Flowers That Don’t Exist

Scammers Sell Seeds for Exotic AI-Generated Flowers That Don’t Exist

Scammers are selling seeds for plants that don’t exist with spectacular, AI-generated images of technicolor leaves that bloom in the shape of birds, butterflies, and cat heads. This type of fake seeds scam predates widespread access to AI image generators, but the ability to easily create these images has made the scam more widespread, especially on big online retailers like eBay, Amazon, and Etsy, which are unable to keep up with the flood of scam plant sellers on their platforms. 

  •  

Wikipedia Cofounder Larry Sanger Banned From Site for ‘Canvassing’

Wikipedia Cofounder Larry Sanger Banned From Site for ‘Canvassing’

Larry Sanger, one of Wikipedia’s cofounders, was banned from editing the site indefinitely after other editors determined he was canvassing, or in other words, calling on his followers off platform in order to influence Wikipedia’s content. 

Sanger has spent more than a decade criticizing Wikipedia for what he claims is an ideological, left-wing bias on a variety of topics, and on X has framed this recent ban as further proof of everything that’s wrong with Wikipedia. The New York Post took that bait and last night published an article with the headline “Left-leaning Wikipedia blocked founder from editing site—after he campaigned to make it more balanced.” 

Wikipedia editors obviously reject that framing and say that Sanger was banned for wielding his followers to sway discussion and decision making on Wikipedia. The discussion that led to the decision to ban Sanger concluded with what an editor called a “clear consensus” to ban Sanger.

“There is general agreement among participants that he has engaged in off-wiki canvassing and is not here to constructively build the encyclopedia,” the editor said in a note closing the discussion. “There is also a significant concern shared by many editors that his actions constitute calls for outing.”

While Sanger has been railing about bias on Wikipedia for years, the specific issue here is around his WikiProject Intellectual Diversity. WikiProjects are group efforts among Wikipedia volunteers to deal with certain issues on the site. For example, in 2024 I wrote about WikiProject AI Cleanup, a group of volunteers who focus on removing AI-generated content from the online encyclopedia. Sanger’s WikiProject Intellectual Diversity, as its name implies, aims to bring more intellectual diversity to the site, mostly meaning more right-leaning perspectives. 

Sanger’s WikiProject Intellectual Diversity and its goals alone do not merit a ban according to Wikipedia’s policies. The problem, according to Wikipedia editors, is that during the discussion about whether to allow WikiProject Intellectual Diversity to become an official WikiProject, Sanger invited his 91,000 followers on X to influence that discussion. 

“Wikipedians are now debating whether my proposed WikiProject Intellectual Diversity should be permitted to become an official WikiProject (club/group of editors),” Sanger said on X on Friday and linked to the Wikipedia talk page about the issue. “Lots opposed. Also lots in favor.”

“Can I still join the movement?” one person replied to Sanger on X

“Let's just say that if I answer that question one way or another, the playground moms who rule Wikipedia might block me,” Sanger responded. 

As one volunteer wrote in the discussion page about whether to ban Sanger:

“Since the return from his self-imposed exile pretty much all he has done is try to start a right-wing/conservative pressure group within Wikipedia not to improve articles on topics that may be under-represented or highlight high-quality sources that could be utilised more, but to instead attempt to rewrite policies and guidelines to his political bent while throwing baseless aspersions about the conduct of many users (mostly those in privileged positions such as admins) and alleging they're being funded by shadow money. Frankly if this was anyone else claiming all this with the way he is, we'd have shown them the door long ago.”

Ilyas Lebleu, another Wikipedia volunteer and admin, told me that they had warned Sanger about similar behavior two months ago, but that Sanger ignored them. 

“Larry tried to frame the community discussion as a pseudo-legalistic process, bringing a list of ‘charges’ and ‘counts’ from ‘prosecutors,’ instead of an open community discussion,” Lebleu said. 

Discussions about potential bans are supposed to remain open for at least 72 hours. While consensus that Sanger had violated Wikipedia policies was clear, Sanger was banned at some point before that deadline. He was then briefly unbanned, and then again indefinitely banned once 72 hours had elapsed and the discussion about the ban closed. 

“Wikipedia has become more of a mob-rule anarchy than ever,” Sanger said in a statement sent to me by a spokesperson. “In the kangaroo court in which a mob ousted me, Wikipedia’s administrators showed that they don’t appear to value details like formal charges, a designated prosecutor, basic decorum, distinction between prosecution and judge, dispassionate adjudication, and so forth. They have no proper system other than triggering a mob to selectively enforce their hodgepodge of vague rules.”

“Now that same mob has blocked me for trying to bring an intellectually diverse group of thinkers and editors to the site,” Sanger continued. “Subscribing to their groupthink is now an official requirement of being a member in good standing. Something must change, and now. I only wonder if the system as it currently stands can even allow the discourse necessary to fix the system.”

Sanger’s claim that Wikipedia has a left-leaning bias isn’t unique or new. Elon Musk has railed against the site for years as well, an effort that culminated with the launch of his highly flawed, AI-generated Grokipedia. But the stakes for Wikipedia as a reliable source of information are higher than ever as every corner of the internet is struggling to deal with a flood of AI-generated, error-filled slop. 

  •  

Judge Rules Blacked.com Can Sue Meta for Scraping Its Porn

Judge Rules Blacked.com Can Sue Meta for Scraping Its Porn

A federal judge has rejected Meta’s attempt to dismiss a lawsuit from Strike 3 Holdings, the company that owns popular sites like Blacked, Vixen, and Tushy, for scraping its porn videos. 

The decision shows Meta’s nonsensical justification for scraping massive amounts of copyrighted material from the internet in order to train its AI models, and is notable for adult content creators, who have been scraped for model training data long before the current generative AI boom.

Strike 3 Holding first filed its lawsuit almost a year ago after internal Meta emails revealed in a different lawsuit showed that the company downloaded over 81 terabytes of data by scraping Anna’s Archive, a massive open search search engine for torrenting copyrighted material including books, movies, TV shows, and porn. A Strike 3 Holding investigation found that 47 IP addresses belonging to Meta were used to torrent 2,396 of its videos a total of 6,008 times between 2018 and 2025. On Thursday, Judge of the United States District Court for the Northern District of California Judge Eumi K. Lee rejected Meta’s attempt to dismiss the lawsuit, allowing it to move forward. 

Meta argued that Strike 3 Holdings failed to show that Meta actually intended to use Strike 3 Holdings’ videos to train its AI models and that Meta, the company, was actually responsible for downloading the videos, as opposed to rogue employees downloading porn on company time from company IP addresses. 

According to the judge’s ruling, Strike 3 Holdings’ investigation showed coordination across Meta’s IP addresses that proved “a coordinated effort to gather data,” as opposed to the action of random employees. Specifically, Strike 3 Holdings showed that Meta’s IP addresses torrented files with similar file names on the same day, ranging from porn to cartoons and sitcoms, suggesting the company was downloading files based on key terms. 

“For example, IP Ranges A and F torrented the following files on December 15, 2022: ‘Teen Sex Sessions 2 (2012),’ ‘Teen Titans Go to the Movies (2018),’ ‘Teens Love Tats XXX,’ ‘TeensLoveAnal.16.09.30.Amara,’ ‘Teenfidelity Pics,’ ‘TeensLoveAnal.16.06.10.Casey,’ ‘Teenage Mutant Ninja Turtles (1987-1996),’ ‘Teen Mom Girls Night In S02E08,’ ‘TeenyTaboo.22.12.07.Kiana,’ and ‘TeenageDelinquents.Maryjane,’” the decision says. “On the same day, a Corporate IP Address was used to torrent ‘TeenCurves.22.12.09.Willow.’ The connection between these files is plain: The word ‘teen’ appears in every file name.”

The judge said that Meta suggesting that its IP addresses downloading all these files at the same time was the work of different individual Meta employees acting independently “strains credulity.”

The judge also explained that whether Meta actually used Strike 3 Holdings’ videos to train its AI models is irrelevant because Meta violated Strike 3 Holdings’s copyright when it torrented its videos. It illegally downloaded the files and also “seeded” them, meaning they distributed the pirated to other users.

“In sum, Plaintiffs [Strike 3 Holdings] have plausibly alleged that Defendant [Meta] is liable for direct, vicarious, and contributory copyright infringement based on the torrenting of their films,” the decision said. “Defendant’s motion to dismiss is therefore DENIED.”

  •  
❌