The practice of destroying books to train AI models has sparked concern among book lovers, who fear that physical copies of rare texts may be lost forever. AI companies often use destructive methods, such as cutting book spines and scanning pages, to quickly train their models on large volumes of text. However, this approach raises ethical concerns, particularly when it comes to preserving fragile or historically significant books. The Internet Archive has long emphasized the importance of careful, non-destructive scanning to protect rare collections.
The Archive’s 2021 post detailed the work of Eliza Zhang, a book scanner who has been with the organization since 2010. Zhang described how she carefully scans fragile books, using a foot pedal to raise the scanner glass and adjusting cameras to ensure readability. She also noted the importance of scanning fold-outs and inserts to avoid losing bonus materials. The process requires patience and attention to detail, as even a single skipped page or blurry image can trigger the Archive’s proprietary software to stop the scan and prompt a retry.
According to the source, the Internet Archive has experimented with automated scanners but found them ineffective for brittle or rare books. The organization’s approach relies on human expertise, with Zhang having scanned over 3 million pages, 14,000 foldouts, and 18,000 items in her career. Chris Freeland, the director of library services, confirmed that Zhang’s method remains the best description of the Archive’s current scanning process.
Source: arstechnica