Prominent U.S. newspapers The Seattle Times and Newsday have filed a copyright infringement lawsuit against OpenAI and Microsoft. The publishers claim that their news content has been used without authorization as training data for generative AI.
This legal action focuses on how large language models (LLMs), including OpenAI's ChatGPT, improperly collect and utilize copyrighted articles as training data. The plaintiffs point out a lack of compensation and legitimate licensing agreements for data utilization in AI development.
While vast amounts of text data are essential for improving the performance of AI models, legal friction is intensifying between media companies and tech firms regarding the copyrights of the source publications. This case follows previous lawsuits filed by other major media corporations, once again highlighting that the handling of intellectual property rights is a critical issue that will determine the sustainability of the AI industry.
The plaintiffs are seeking damages while also advocating for stronger copyright protection. The verdict in this trial could have a decisive impact on how data is collected in future generative AI development and on the establishment of licensing frameworks between the publishing industry and AI companies.