The controversy over the legality of training data for generative artificial intelligence has recently resurfaced. According to unsealed documents from The New York Times' copyright lawsuit, a senior research director at Microsoft stated in internal discussions that the current training methods for large language models may constitute the 'largest-scale labor theft in human history,' a comment that has drawn industry attention.
