An article in The New Yorker delved into the Panama Project, an initiative by AI firm Anthropic to train its Claude model. Under the pretense of a secretive entity, Anthropic acquired a vast array of obscure academic texts from the 1970s and beyond, along with other resources. These materials were then dispatched to industrial sites, where they were shredded, scanned, and ultimately the physical copies discarded—all in a bid to amass data from books across the globe.
From a legal standpoint, a judge determined that utilizing legally procured books to train AI and transforming their formats falls under fair use. Anthropic has also reached settlements in related copyright disputes. Yet, the crux of the controversy lies in transparency and cultural implications: Anthropic's so-called 'research library' serves exclusively as a data mine, shielding from public view which books and information sources have been assimilated by the model. Experts highlight that AI reduces books to malleable data, stripping them of their historical backdrop, narrative architecture, reading sequence, and other intrinsic worth. This clandestine obliteration is deemed more alarming than the physical annihilation of books.
Nonetheless, some contend that obscure books integrated into models might attain a heightened value, and there will always be individuals yearning for the authentic reading encounter. Moreover, Anthropic has delayed the release of its prospectus, and at present, there is a lack of publicly available information regarding analogous practices among Chinese large-scale model companies.
