Loading the Elevenlabs Text to Speech AudioNative Player...

A new lawsuit filed against Meta is intensifying the legal battle over how AI companies train their models, with publishers accusing the company of deliberately using pirated books and articles to build its Llama AI system instead of licensing them legally.

The lawsuit was brought by publishers Hachette, Macmillan, McGraw-Hill, Elsevier, and Cengage alongside author Scott Turow, who allege that Meta trained Llama using more than 267 terabytes of copyrighted materials sourced from shadow libraries, including the piracy database LibGen. According to the complaint, the amount of material involved exceeds the entire print collection of the Library of Congress.

Plaintiffs described the alleged conduct as “one of the most massive infringements of copyrighted materials in history.”

What makes this case particularly significant is not just the scale of the alleged infringement, but the paper trail cited throughout the lawsuit. According to the filing, Meta internally considered spending as much as $200 million to license books and research material for AI training before ultimately deciding against it. Plaintiffs claim the discussion was eventually escalated to Meta CEO Mark Zuckerberg, who allegedly rejected the licensing approach.

One internal employee message cited in the complaint stated: “If we license one single book, we won’t be able to lean into the fair use strategy.”

Subscribe for free to continue reading this article

Subscribe Subscribe

Already Have an Account? Log In