When Large Models ‘Devour’ Internet-Wide Data: How to Define the Boundary Between ‘Fair Use’ and Infringement?
7 hour ago / Read about 0 minute
Author:小编   

As large model companies ramp up investments in computing infrastructure, copyright issues concerning AI training data have emerged as a growing, yet often overlooked, concern. In mid-August 2026, news broke that Apple intended to spend hundreds of millions of dollars to acquire news datasets to enhance its AI-powered Siri, sparking widespread industry debate. This development followed a wave of global copyright lawsuits related to AI training data: In December 2023, The New York Times sued OpenAI and Microsoft for unauthorized use of millions of its articles to train models like ChatGPT, claiming billions in damages. In June 2026, nearly 400 U.S. print media outlets jointly sued OpenAI and Microsoft, accusing them of systematically scraping news content for AI product training. In the literary sphere, six authors—including two-time Pulitzer Prize winner John Carreyrou—filed a collective lawsuit against Anthropic, OpenAI, and others for “intentional theft.” Domestically, in November 2025, Shanghai’s first AI large model copyright infringement case was adjudicated, centering on disputes over the reproduction rights of animated characters from Battle Through the Heavens.

Apple’s proposed “pay-per-use” model, while hailed by some legal experts as a potential breakthrough in AI-era content licensing, remains legally contentious. Fu Gang, a partner at Beijing Dacheng Law Firm, emphasized that copyright law prioritizes whether explicit authorization exists for the actual use, rather than the payment structure. Infringement risks persist if the licensing scope fails to cover the specific usage scenario. Technically, metrics for pay-per-use remain ambiguous, with unresolved questions such as how to calculate usage when multiple news articles are synthesized into a single search summary, or how to account for repeated API calls due to technical errors.

From a market perspective, news of Apple’s news dataset procurement triggered a rally in China’s A-share “AI dataset” stocks, with companies like Dooke Culture and CITIC Press following suit. Gou Yurui, an analyst at Southwest Securities, argued that AI is driving the media content industry to transition from a “consumer goods” model to dual pricing for “data + IP assets.” However, he cautioned that the implementation of a token-based valuation system hinges on observing licensing transactions and substantive policy advancements.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic