If you can truly appreciate an old book—and maybe even marvel at how its fragile, yellowing pages contain some of the earliest ways that people tried to make sense of the world around them—then headlines about tech companies that are destroying books to train AI likely torture a tender part of your soul.
It’s indeed depressing to imagine piles of book spines waiting to be fed into wood chippers while torn-out pages are cropped, scanned, and trashed. But that’s the cheapest and easiest way to scan books as fast as possible, and AI companies are in a race to advance their models by training on the kind of engaging, high-quality long-form texts that can only be found in books. So book lovers fear it’s likely that the practice is happening on a grander scale than is currently being reported and that some physical copies of books will be lost forever.
What makes this destruction extra painful, though, is that it doesn’t have to be this way. //
“At the Internet Archive, this is how we digitize a book,” the tweet said. “We never destroy a book by cutting off its binding. Instead, we digitize it the hard way—one page at a time.” //
“The job requires keen concentration,” Zhang said, since the pages of “very old, fragile books” are “paper thin.” In the post, Andrea Mills, who helps lead the Archive’s book-scanning operations, explained that “clean, dry human hands are the best way to turn pages.” //
The biggest fear for people who want to see books preserved through the training process is that AI firms will callously pulp rare books that can never be replaced.