OpenAI scored a major success in a recent case in which Raw Story Media, Inc. and AlterNet Media, Inc lost a motion to dismiss their case alleging that OpenAI’s removal of copyright management information (CMI) from thousands of articles prior to using them to train its ChatGPT product violated Section 1202(b)(i) of the DMCA. In dismissing the case, the court relied on the U.S. requirement to establish Article III Standing (that the injury must be “concrete and particularized” and “actual or imminent”). While the standing requirement is well known in U.S. law and has been litigated in other copyright cases, this is the first time it has been used to dismiss a breach of the DMCA’s CMI provisions. As such, the decision in Raw Story Media v OpenAI No. 24 Civ. 01514.(S.D.N.Y Nov 7, 2024) is not only important in relation to whether the removal of CMI to train large language models (LLMs) can give rise to a cognizable claim, but also more generally to the right of copyright owners to pursue other claims for violation of the DMCA CMI provisions. The decision is also very important as the plaintiffs in the case are alleging only violations of the DMCA CMI rights to avoid the more difficult claims of alleging that the training and outputs of large language models infringed the U.S. Copyright Act’s exclusive rights and the defenses being raised in such cases.
Background to the Case
The plaintiffs are news organizations that collectively publish hundreds of thousands of articles and opinion pieces. They allege that OpenAI’s stripped training datasets of their copyrighted works of the accompanying CMI, such as author and title information before using them for training ChatGPT. According to the plaintiffs, these datasets, including WebText, WebText2, and Common Crawl, were instrumental in training ChatGPT.
The plaintiffs argued that the removal of the CMI was a violation of Section 1202(b)(i) of the DMCA, which prohibits the intentional removal or alteration of CMI knowing it will facilitate copyright infringement. Section 1202(b)(i) of the DMCA provides that:
No person shall, without the authority of the copyright owner or the law, … intentionally remove or alter any [CMI] … knowing, or with respect to civil remedies under section 1203, having reasonable grounds to know that it will induce enable, facilitate or conceal an infringement of any right under this title.
This section of the DMCA has been found in the Ninth Circuit to incorporate a so-called double-scienter requirement: that the defendant both know “that CMI has been removed or altered without authority of the copyright owner or the law” and know, “or hav[e] reasonable grounds to know that such distribution will induce, enable, facilitate, or conceal an infringement.”
The plaintiffs sought both injunctive relief and damages. They did not make infringement claims that using their articles to train ChatGPT unlawfully reproduced their content or that ChatGPT’s output infringed their copyrights. In this respect, the plaintiff’s claims eschewed the much broader claims of infringement brought against OpenAI, Google, Nvidia, Meta, and Anthropic in other cases where claims of infringement were alleged in both the training and output of their LLMs.
Court’s Holding
OpenAI brought a motion to dismiss the plaintiffs’ complaint arguing lack of Article III standing to assert their claims. Whether a claimant has standing is the threshold question in every federal case, determining the power of the court to entertain the suit.
The plaintiffs argued that they had standing to pursue their claims for two essential reasons. First, because the “the unlawful removal of CMI from a copyrighted work is a concrete injury”. Second, because there is a substantial risk that ChatGPT will “provide responses to users that incorporate[] material from Plaintiffs’ copyright-protected works or regurgitate[] copyright-protected works verbatim or nearly verbatim.”
The court sided with OpenAI which found that the claim failed to meet the standing requirement for a damage claim because the plaintiffs did not allege any actual adverse effects stemming from the alleged DMCA violation.
OpenAI also argued that plaintiffs lacked standing to seek injunctive relief because they failed to allege facts to show that the risk of ChatGPT reproducing Plaintiffs’ work, in whole or in part, absent the requisite CMI was “substantial.” The court agreed, based on its opinion that it is unlikely the output generated from the more current version of ChatGPT would contain infringing reproductions of the plaintiffs’ works.
I agree with Defendants. Plaintiffs allege that ChatGPT has been trained on “a scrape of most of the internet,” which includes massive amounts of information from innumerable sources on almost any given subject. Plaintiffs have nowhere alleged that the information in their articles is copyrighted, nor could they do so. When a user inputs a question into ChatGPT, ChatGPT synthesizes the relevant information in its repository into an answer. Given the quantity of information contained in the repository, the likelihood that ChatGPT would output plagiarized content from one of Plaintiffs’ articles seems remote. And while Plaintiffs provide third-party statistics indicating that an earlier version of ChatGPT generated responses containing significant amounts of plagiarized content, Plaintiffs have not plausibly alleged that there is a “substantial risk” that the current version of ChatGPT will generate a response plagiarizing one of Plaintiffs’ articles.
The court dismissed the claim allowing the plaintiffs to seek leave to file an amended Complaint. However, the court was skeptical as to whether the Complaint could be amended to state a cognizable claim because the gravaman of the plaintiffs’ complaint was the infringement of their copyrights in the training of ChatGPT and not the removal of CMI from the works.
Let us be clear about what is really at stake here. The alleged injury for which Plaintiffs truly seek redress is not the exclusion of CMI from Defendants’ training sets, but rather Defendants’ use of Plaintiffs’ articles to develop ChatGPT without compensation to Plaintiffs. (“The OpenAI Defendants have acknowledged that use of copyright-protected works to train ChatGPT requires a license to that content, and in some instances, have entered licensing agreements with large copyright owners… They are also in licensing talks with other copyright owners in the news industry, but have offered no compensation to Plaintiffs.”). Whether or not that type of injury satisfies the injury-in-fact requirement, it is not the type of harm that has been “elevated” by Section 1202(b)(i) of the DMCA. See Spokeo, 578 U.S. at 341 (Congress may “elevate to the status of legally cognizable injuries, de facto injuries that were previously inadequate in law.”). Whether there is another statute or legal theory that does elevate this type of harm remains to be seen. But that question is not before the Court today.
Comments
The Raw Story v OpenAI case appears to reflect a tactical litigation strategy to allege in their Complaint only a violation of the DMCA CMI rights without also alleging that the training or output of LLMs infringes the reproduction and other exclusive copyright rights and to try and avoid the fair use defense to such claims. This strategy was implicitly recognized by the court that observed that the alleged real complaint of plaintiffs were the infringements of copyright associated with training the ChatGPT LLM. It will be interesting to see whether the plaintiffs can get leave to amend their Complaint to pass muster under the standing requirement without also expanding the claims to the more common claims that would be clearly subject to the fair dealing and other potential defenses. As such, what happens in this case could have substantial repercussions in other claims made against providers of LLMs.
There are now over 30 cases in the U.S. alleging copyright infringement in the training and uses of LLMs. (This is in addition to cases in other countries including the recent case by GEMA in Germany). Some of these cases also allege violations of the CMI provisions in the DMCA and cases have been dismissed on the grounds that the LLMs’ outputs are not identical copies of the works from which the CMI has been removed. See, Barry Sookman, Generative AI litigation: the Github and Tremblay decisions; Barry Sookman, AI models and copyright infringement, Andersen v. Stability AI. The question of law as to whether violation of the CMI provisions must be on identical copies content has been appealed to the U.S. Ninth Circuit Court of Appeals in the Github case.
The Article III standing issue has been raised in other copyright cases so it is not a novel defense. [i] It has also been raised under the DMCA such as in relation to standing to sue for misrepresentations in takedown notices.[ii]
The standing issue is also being used by NVIDIA in an attempt to dismiss David Millette’s complaint alleging state law claims related to NVIDIA’s scraping videos to train NVIDIA’s AI model.
This is the first case, however, to consider whether the removal of CMI without any allegation that the training or outputs infringe copyright present a cognizable claim. While a very significant case on the question of DMCA claims for violation of the CMI provisions of the DMCA, the decision may be even more important if it leads to further standing defences to other claims of infringement in future AI infringement cases or to CMI claims that bypass the more common claims and defenses to training LLM models.[iii]
[i] See, for example, Naruto v. Slater, 888 F.3d 418 (9th.Cir. 2018); Infogroup Inc. v. Office Depot, Inc. 688 F.Supp.3d 1179 (S.D.Fla. 2023); Michael Pellis Architecture PLC v. M.L. Bell Construction LLC, F.Supp.3d 594 (E.D. Vir. 2023).
[ii] See, Handshoe v. Perret 270 F.Supp.3d 915 (S.D. Miss.2017)
[iii] As to whether fair use applies to DMCA CMI claims, see, Lim, A Survey of the DMCA’s Copyright Management Information Protections: The DMCA’s CMI Landscape after All Headline News and McClatchey, 6 WASH. J. L. TECH. & ARTS 297 (2011). Available at: https://digitalcommons.law.uw.edu/wjlta/vol6/iss4/5