Key Points
- The Japan Book Publishers Association has formally questioned major distributor Nippan over allegations that vast quantities of books were sold to US AI company Anthropic for training without author o
- The inquiry follows US copyright court records and sightings of bulk secondhand book sweeps, with reports revealing that over 50 tons of books—roughly 100,000 volumes—were exported abroad.
- Publishers strongly condemned the practice, arguing that physical cultural works and specialized literature should never be funneled into AI training datasets without proper authorization.
Summary A major controversy has erupted in the East Asian publishing world after allegations surfaced that massive quantities of Japanese books were systematically acquired and shipped abroad to train generative artificial intelligence models. The Japan Book Publishers Association, a prominent trade organization representing roughly 380 publishing houses including major literary giants such as Kodansha and Bungeishunju, has officially demanded answers from Japan Publications Distribution (Nippan), one of the country's primary wholesale book distributors.
The controversy stems from court filings in an ongoing copyright infringement lawsuit filed in the United States against AI developer Anthropic. According to reports cited by Kyodo News, the legal records detailed a pipeline through which books were procured from Japanese distributors to build AI training datasets. Suspicions had already been mounting across Japan in recent months when secondhand bookshops reported unprecedented bulk purchases of academic and specialized volumes covering subjects like philosophy, history, medicine, and law. Many in the industry suspected that AI firms were purchasing physical copies en masse simply to scan their contents into digital datasets before discarding the physical copies.
Broadcast network Nippon TV revealed that a subsidiary of a leading Japanese book distributor had exported more than 50 tons of books to the United States—an enormous volume estimated at around 100,000 titles. Officials from the Japan Book Publishers Association have voiced strong objections, emphasizing that it is deeply regrettable if books were sold with the foreknowledge that they would be fed into AI training pipelines. The association stated that literature and copyrighted academic works must never be sold for AI ingestion without the explicit consent of publishers and authors.
As AI developers race to acquire high-quality, linguistically diverse training data, the incident underscores the growing international tension between creative rights holders and tech companies. While the article was reported in South Korea due to shared regional concerns over copyright protection and intellectual property exploitation by global tech giants, it highlights an unprecedented flashpoint in how physical cultural works are digitized and harvested in the generative AI era.
Sponsored · Ad