On October 15, Salesforce found itself embroiled in a class-action lawsuit initiated by two novelists. The plaintiffs alleged that Salesforce's xGen series of large language models were trained on pirated books. Specifically, they accused Salesforce of unauthorized use of copyrighted book datasets in the development of its AI models. This situation echoes a prior case where generative AI firm Anthropic settled similar infringement claims for a staggering $1.5 billion. It is anticipated that the Salesforce case will also reach a settlement.
This legal action could cast a shadow of doubt over the reliability of Salesforce's models and the legitimacy of its training datasets in the eyes of corporate clients. Consequently, when vetting AI suppliers, companies must conduct thorough due diligence on the origins of training data and the terms of compensation for any potential infringements. Such lawsuits have the potential to impede the broader adoption of AI technology, as they highlight the legal and ethical complexities surrounding data usage in AI development.
