A new lawsuit has been levied against Meta, alleging that CEO Mark Zuckerberg authorized the use of pirated datasets for training its Llama AI models? The case, Kadrey v? Meta, includes claims from prominent authors such as Sarah Silverman and Ta-Nehisi Coates? The plaintiffs argue that Meta trained its AI on copyrighted materials from LibGen, a known source of pirated books and articles?
Allegations Against Meta
New court documents reveal that Zuckerberg reportedly approved Meta’s use of the LibGen dataset for AI development, despite knowing its pirated nature? Meta employees have internally referred to LibGen as a problematic dataset that could endanger Meta’s dealings with regulators? Testimonies from Meta suggest an attempt to sidestep traditional licensing negotia�tions by citing fair use as a defense, a stance many creators oppose?
Details of the Case
Evidence shows Meta took measures to obscure its use of LibGen data? Meta engineer Nikolay Bashlykov allegedly developed a script to remove copyright information from the dataset? This action appears to have been an attempt to conceal the supposed infringement? Further, the plaintiffs argue that Meta also used torrenting methods to obtain data, potentially violating copyright laws by redistributing content without permission?
Implications for Meta
The lawsuit illuminates significant allegations against Meta, including claims that the company tried to erase copyright metadata to cover its tracks? Judge Vince Chhabria, overseeing the case, has criticized Meta�s efforts to keep details under wraps, suggesting that such actions aim to prevent reputational damage rather than protect business secrets?
Meta insists on relying on the fair use doctrine, hoping to dismiss the claims as it did previously in 2023 when other AI-related copyright claims were unsuccessfully brought against them? The court�s future rul�ings will determine whether Meta’s defenses hold up or if the authors’ claims lead to repercussions?
The lawsuit focuses on activities related to Meta’s initial Llama models, not the later versions, leaving the lawsuit’s impact on the company’s AI strategy uncertain? A verdict in favor of the plaintiffs could set a significant precedent for how AI models are trained on copyrighted content?
TechCrunch reached out to Meta for comment, but the company has yet to provide a statement? The case�s developments will likely influence both the tech industry and future regulatory perspectives on AI model training?

