Save rare books before AI shreds them

AI firms are pulping physical books for training data. Here’s how libraries and collectors can digitize rare works before they disappear forever.

4 min read

Save rare books before AI shreds them cover

Last month, a friend who runs a small university archive showed me a pallet of 19th-century medical texts. They were headed for a recycling plant, not because the books were damaged, but because an AI company had bought the collection to shred for training data. The irony stung: these were the last surviving copies of some titles, now slated for destruction in the name of progress.

Why rare books are vanishing

AI firms need vast amounts of text, and they’re not picky about the source. Physical books, especially older ones, are cheap and plentiful. Many are being acquired in bulk from estate sales, library discards, and defunct collections. Once scanned, the books are often pulped to avoid storage costs. The problem isn’t just the loss of the physical object, it’s that many of these books exist nowhere else. Digitization by AI companies is usually low-resolution, unsearchable, and locked behind proprietary systems. When the books are gone, so is the chance to preserve them properly.

What makes a book worth saving

Not every old book is rare, but many are. Here’s how to spot the ones that need urgent digitization.

  • First editions of historically significant works, especially those with marginalia or inscriptions.

  • Books published before 1928 (U.S. public domain cutoff) that aren’t already widely digitized.

  • Local or regional publications, like city directories, church records, or small-town newspapers.

  • Works with unique physical attributes: hand-colored maps, fold-out illustrations, or unusual bindings.

  • Books from defunct publishers or authors whose other works have been lost to time.

How to scan books without destroying them

You don’t need a million-dollar setup to digitize books well. Here’s what actually works.

  • Use a flatbed scanner for fragile or oversized books. Overhead scanners like the CZUR ET16 work well for most materials and don’t require pressing the book flat.

  • For large collections, consider a book cradle with a DSLR camera. A basic setup costs under $500 and can capture 300 pages per hour.

  • Scan at 600 DPI for text, 1200 DPI for illustrations. Save files as TIFF for archival quality, then convert to PDF or DjVu for distribution.

  • Use OCR software like Tesseract or ABBYY FineReader to make the text searchable. This is critical for research use.

  • Store files in multiple locations: a local hard drive, a cloud backup, and an institutional repository if possible.

Where to share digitized books

Digitization is useless if no one can access the files. Here’s where to upload them for maximum impact.

  • Internet Archive: The gold standard for public-domain books. They handle hosting, OCR, and long-term preservation.

  • HathiTrust: A partnership of academic libraries. Great for scholarly works, but requires institutional affiliation for some features.

  • Project Gutenberg: Focuses on plain-text versions of public-domain books. Low-tech but highly accessible.

  • Local digital libraries: Many cities and universities have their own repositories. These are ideal for region-specific materials.

  • Specialized collections: Sites like the Medical Heritage Library or the Biodiversity Heritage Library focus on niche topics.

The legal risks and how to avoid them

Copyright law is a minefield, but there are safe paths. In the U.S., anything published before 1928 is public domain. For newer books, check the copyright status using the Stanford Copyright Renewal Database or the U.S. Copyright Office’s records. If a book is still under copyright, you can still digitize it for preservation purposes under fair use, but you can’t distribute it publicly without permission. Libraries and archives have more leeway here, thanks to Section 108 of the Copyright Act.

What you can do today

You don’t need to be a librarian to make a difference. Start with these steps.

  • Inventory your local library’s discard pile. Many libraries sell or recycle books they no longer need. Ask if you can review them before they’re tossed.

  • Partner with historical societies. They often have collections they can’t afford to digitize. Offer to scan a few books in exchange for access.

  • Run a community scanning day. Set up a station at a local event and invite people to bring their rare books. You’ll be surprised what turns up.

  • Donate to digitization projects. Organizations like the Internet Archive and Project Gutenberg rely on volunteers and funding to keep going.

  • Advocate for better policies. Push for laws that require AI companies to disclose what books they’re using for training data, and to preserve the physical copies.

The books being shredded today won’t be around tomorrow. Every scan you make is a hedge against that loss. Start small, but start now.

Building something with AI? Let's talk.

I design and ship production AI and full-stack products for US teams. See how I can help.

View all services

Join the newsletter

Be the first to read our articles.