Today, the Ministry of State Security released a report highlighting the presence of mixed quality issues within AI training data. These issues encompass false information, fictional content, and biased viewpoints, collectively leading to data source pollution and posing significant new challenges to AI security. Research conducted indicates that a mere 0.01% contamination of training data with false texts can result in an alarming 11.2% increase in the model's output of harmful content.
