How Meta Platforms faces copyright suit against regarding AI Training techniques

Prominent writers, including Michael Chabon and Sarah Silverman, are suing Meta Platforms, the parent corporation of Facebook and Instagram, alleging copyright infringement. The main charge is against Meta for allegedly using thousands of copyrighted books without authorization in order to develop Llama, their artificial intelligence language model.The company apparently moved forward with the controversial dataset even after Meta's legal team sent out strong warnings about the potential legal ramifications of using stolen books for AI training. When evidence from conversation logs emerged, indicating researcher Tim Dettmers of Meta discussing the dataset's procurement in a Discord server, the matter became more complicated.The chat logs show that Dettmers corresponded with Meta's legal department, raising questions on the permissibility of using book files as training data. The legal staff recommended against using right away.Citing concerns about "books with active copyrights," the legal team recommended against using the material straight away. Participants in the discussion discussed whether training using such data may be permitted by the fair use doctrine, a legal theory in the United States that shields some unapproved uses of works that are protected by copyright.The case, which was first started in the summer, recently combined two unrelated lawsuits against Meta. A California judge dismissed a portion of the Silverman complaint last month, which prompted the authors to revise their allegations and indicated a changing legal landscape.Beyond Meta, the ramifications of this legal dispute could have an impact on the larger AI sector. If these cases are successful, it might become more expensive to create AI models with a lot of data, which would expose businesses to more scrutiny and requests for payment from content providers. Furthermore, AI businesses like Meta may be forced by new European regulations to reveal the data that was used to train their models, putting them at further legal danger.Meta's Llama models, particularly the most recent model, Llama 2, which was published in the summer, are at the center of the dispute. Although Llama 2, a possible disruptor in the generative AI software market, was trained using the "Books3 section of ThePile" in the first edition, Meta has not released information regarding the training data for this version

Enjoyed this article? Stay informed by joining our newsletter!

Comments

You must be logged in to post a comment.

About Author