Amazon, which began as an online bookseller, is reportedly using rare books to train its AI models. According to 404 Media, the company is purchasing large quantities of rare texts, removing their spines, and scanning them for AI training data. The tracking device placed in a rare book revealed its arrival at an Amazon facility in Las Vegas, known as VGT3. The facility is identified by a symbol of a dinosaur holding a book in its claws. Amazon stated it purchases books through commercial channels to improve the products and services customers use. The company’s move comes as it seeks new sources of training data, especially for large language models (LLMs) that have already ingested much of the available internet content. Rare books, particularly those out of print or unavailable online, offer a unique opportunity for training data. These texts are especially valuable because they were published before 2022, ensuring they were not written by AI-generated text. Training LLMs on AI-generated content risks model collapse, a phenomenon where the quality of an LLM’s outputs degrades after ingesting too much AI-generated text. This practice raises ethical and preservation concerns, as it may lead to the loss of rare and historical texts. The issue highlights the growing demand for training data in the AI industry and the potential consequences of sourcing it from unconventional places.
Source: techcrunch