“The complaint cited GPT-4 as stating that its ‘training data gave it enough exposure to wikiHow content to be able to generalize their visual patterns and formats effectively.’”
On August 21, wikiHow, Inc. filed a complaint in the U.S. District Court for the Southern District of New York against OpenAI, Inc. and eight affiliated entities, alleging violations of the Copyright Act and the Digital Millennium Copyright Act (DMCA). The lawsuit centers on claims that OpenAI copied wikiHow’s how-to articles without authorization to train ChatGPT and to ground its outputs, then used that copied content to generate substitutes for wikiHow’s website.
wikiHow described itself as a publisher of expert-verified instructional articles that has served billions of readers since its founding in 2005. The complaint stated that wikiHow hosts more than 500,000 articles and, at its peak, drew more than 150 million visitors a month. wikiHow owns 1,211 registered copyrights covering 11,211 articles at issue in the case, identified in an exhibit attached to the complaint.
According to the complaint, OpenAI has acknowledged training its large language models (LLMs) “on vast amounts of data from the internet written by humans.” wikiHow alleged that its articles were swept into that training data through crawling and scraping of its websites, whether directly or through intermediary datasets such as Common Crawl, WebText, and WebText2.
wikiHow claimed that OpenAI copied its articles at scale to train the large language models underlying ChatGPT. It further claimed that OpenAI separately retrieves, copies, and uses wikiHow’s articles through retrieval-augmented generation (RAG) systems, which supplement the models’ existing knowledge with external content at runtime. Lastly, wikiHow claimed that ChatGPT generates outputs that reproduce its protected expression, including full or partial verbatim text and its distinctive selection and arrangement of instructional material.
To support its training and output claims, wikiHow pointed to a series of prompts submitted to GPT-4 in late 2025 designed to elicit verbatim text from specific wikiHow articles. In one example, when asked to extract verbatim text from wikiHow’s article on setting boundaries with a mother-in-law, GPT-4 reproduced multiple sentences matching wikiHow’s registered article word for word. Similar outputs were generated for wikiHow articles on the meaning of wedding dreams, differences between individuals born in March and April with the Aries zodiac sign, and the concept of soul ties. wikiHow also alleged that GPT-4 could identify a copyrighted wikiHow image by name, explaining that it recognized the file based on wikiHow’s naming conventions and that its “training data gave it enough exposure to wikiHow content to be able to generalize their visual patterns and formats effectively.”
The complaint alleged that wikiHow’s website logs document extensive and continuing scraping activity. According to the filing, wikiHow’s website received more than 185,000 visits from OpenAI’s published IP addresses between March and May 2025, and OpenAI’s crawlers reached wikiHow’s site 148,529 times between May and July 2026 alone. wikiHow stated that it implemented robots.txt directives disallowing OpenAI’s GPTBot crawler in August 2023, and later added directives blocking OpenAI’s OAI-SearchBot and ChatGPT-User crawlers after learning of their existence in January 2025. wikiHow alleged that OpenAI’s bots have continued to access its site notwithstanding those technical prohibitions.
wikiHow argued that OpenAI cannot rely on a fair use defense because its use of wikiHow’s content is commercial and involves copying the articles in their entirety rather than in excerpts. It further argued that ChatGPT’s outputs supplant rather than transform wikiHow’s original expression. The complaint also stated that wikiHow approached OpenAI about a licensing arrangement in spring 2024 and exchanged letters with the company, but “OpenAI never seriously pursued a license” despite entering into similar licensing agreements with other publishers, including News Corp.
wikiHow also alleged that OpenAI intentionally removed copyright management information (CMI), such as article titles, author bylines, publication dates, and terms-of-use notices, from every wikiHow article it copied. wikiHow attributed this removal to content-extraction tools, including Dragnet and Newspaper, which separate an article’s body text from surrounding webpage elements such as footers and copyright notices. It argued that, since OpenAI scraped its websites indiscriminately and processed everything it collected using the same extraction tools, the CMI was stripped from every wikiHow article that entered OpenAI’s datasets.
wikiHow asserted four claims, including direct copyright infringement for the copying used in training and RAG systems and the copying reflected in ChatGPT’s outputs, vicarious copyright infringement against several OpenAI entities, and removal of CMI under Section 1202(b)(1) of the DMCA.
Ultimately, wikiHow is seeking actual damages and OpenAI’s profits attributable to the alleged infringement, or in the alternative, statutory damages under Section 504(c), including enhanced statutory damages of up to $150,000 per work for willful infringement. It is also pursuing statutory damages of between $2,500 and $25,000 per violation under Section 1203(c)(3)(B), and a permanent injunction barring OpenAI from operating any model trained on its copyrighted works without authorization.

Join the Discussion
No comments yet. Add my comment.
Add Comment