The legality of training data therefore depends on the existing fair use categories, on licensing, and on contractual arrangements. Three levels:
Copyright. Acquiring and using the material can infringe in itself. In the Medusa case, the court held that capturing images of anime characters to train a model, and then releasing that model, infringed both the reproduction right and the right of communication through information networks. In the Joy of Life "one-click video" case, cutting the series into clips and loading them into a database for AI retrieval led to an award of RMB 800,000 (approx. USD 119,000). Scraping raises the separate question of whether it breaches a site's terms of use or amounts to unfair competition.
Regulation. Article 7 of the Interim Measures for the Management of Generative Artificial Intelligence Services requires that training data and the foundation model come from lawful sources, that they not infringe intellectual property, and that consent be obtained where personal information is involved. Output must also be labeled under the Measures for Labeling AI-Generated and Synthetic Content.
Data. Where the training set contains personal information or important data, the personal information protection regime and the cross-border data transfer rules apply on top.
In practice: keep a provenance log for every source, retain the licensing chain end to end, filter and de-identify high-risk material, and allocate liability expressly in your service agreements.